The AI price war is moving into the cache

For long-running agents, the price of reusing context can matter more than the headline input-token rate.

Comparison of cached-token prices for Claude Opus 5.5 and GPT-6.1 Sol
Visual · AIpreneur editorial graphic

The cheapest token may be one already read

Anthropic lists Claude Opus 5.5 at $4 per million input tokens, $20 per million output tokens and $0.20 per million cache-read tokens. It describes that cache-read price as 60% lower than Opus 5. OpenAI lists GPT-6.1 Sol at $2 per million input tokens, $10 per million output tokens and $0.10 per million cached input tokens.

A simple comparison makes Sol cheaper on each published unit. Yet the operating cost of an agent depends on the mix: fresh context, reusable context, generated output, tool calls and retries. A rate card cannot tell a builder how many attempts a task will need or how much context the system can reliably reuse.

Why agents change the calculation

Long-running coding and research agents repeatedly consult instructions, repositories, documents and conversation history. Anthropic says cache reads make up the majority of costs in agentic and coding workloads. When that pattern holds, a small change in the cache rate can affect the bill more than a larger change in standard input pricing.

Caching only helps when content remains reusable. Constantly changing prompts, poor cache boundaries or frequent invalidation can return the workload to full input prices. Teams need traces from their own tasks rather than assumptions based on a benchmark prompt.

Price the outcome

For a product team, the useful unit is the completed job: a resolved support case, a reviewed contract, a merged change or a finished report. Track the median and tail cost of that job, including failures and human review. Then decide which model and cache policy delivers acceptable quality at a sustainable margin.

The cache price war improves the economics of persistent context. It also makes vague token comparisons less useful. Customers buy an outcome. Builders need to understand the machinery well enough to price that outcome with room for variance.

Price the completed task after measuring cache behaviour, retries and output — not from a model rate card alone.

Sources and review

Developed from an approved AIpreneur post and reviewed against the cited sources on 4 October 2026.

The question behind every AIpreneur piece: why does this matter to someone building, creating or contributing to the AI economy?

← Back to News

Keep exploring