Claude Opus 5.5 pricing: $4/$20 and the cache-read edge

Released September 22, 2026, Opus 5.5 undercuts its predecessor by 20% on the sticker — and by around 40% on typical workloads, per Anthropic's own measurement. The difference is the cache-read rate: $0.20 per million tokens.

prices verified September 26, 2026

The headline numbers

Anthropic's listed prices for Claude Opus 5.5, verified September 26, 2026:

Token typePrice / 1MWhat it is
Input (fresh)$4.00New tokens never seen before
Input (cache read)$0.20Repeated context — 95% discount
Output$20.00Generated tokens, incl. thinking

That $0.20 cache-read line is the whole story of this model. Everything else — the $4/$20 headline, the release timing, the agent positioning — is context for why cache reads matter so much here.

20% off the sticker, ~40% off typical bills

Opus 5.5's list price sits 20% below Opus 5 ($5/$25). Anthropic says the combination of lower rates and lower token use reduces typical workload cost by around 40% versus Opus 5 — a vendor measurement, not a guarantee for every workload. Two factors drive it beyond the raw discount:

  • Lower token use in real testing. GitHub's chief product officer said Opus 5.5 used among the fewest tokens and steps in GitHub's Copilot CLI and VS Code testing — solving more terminal tasks than Opus 5 while using less than half as many steps — so the same task costs fewer billable tokens.
  • Cheaper cache reads compounding. At $0.20/M — down from Opus 5's higher cache rate — every repeated prompt segment costs almost nothing. In workloads with stable system prompts, the cached portion of the bill collapses.

This is exactly the "token price isn't task price" effect: a 20% sticker cut becomes a 40% bill cut because the cheaper model also consumes fewer billable tokens per task.

Cache-read economics: where agents win

Prompt caching works like this: when you resend context the model has already seen — system prompts, tool definitions, conversation history — the provider skips reprocessing it and bills those tokens at the cache-read rate instead of the fresh-input rate. On Opus 5.5 that means $0.20/M instead of $4.00/M: a 95% discount.

Now consider a coding agent's actual loop. Each step of the loop re-sends the full conversation: the repo context, the system prompt, the tool schemas. In production agent deployments, cache-hit rates above 50% are common — often 70–90% once the context stabilizes.

A worked example

A coding agent session burns 2M input tokens total, with 80% cache hits, and generates 200K output tokens:

Fresh input: 0.4M × $4.00 = $1.60
Cached input: 1.6M × $0.20 = $0.32
Output: 0.2M × $20.00 = $4.00
Total: $5.92 for the session

Without caching, the input line alone would be $8.00 instead of $1.92. The cache turns a $12 session into a $5.92 session. This is why Opus 5.5 is positioned as the agent-economics pick: the longer the agent runs, the larger the cached share, and the more the effective price approaches the cache-read rate rather than the headline rate.

Design rule: structure agent prompts to maximize repeat context. Put stable instructions, tool schemas, and repo overviews first; put the changing query last. Anthropic's cache keys on prefix, so stable-prefix design is what earns the $0.20 rate.

Opus 5.5 vs the field

ModelInput / 1MOutput / 1MCache read / 1M
Claude Opus 5.5$4.00$20.00$0.20
GPT-6 Astra$10.00$50.00discounted*
GPT-6 Sol$2.00$10.00discounted*
Claude Fable 5.1$10.00$50.00$0.25
Grok 4.7$2.00$6.00$0.50

*OpenAI discounts cached input on the GPT-6 family; check current docs for exact rates.

The comparison that matters: Opus 5.5 vs GPT-6 Astra at $10/$50. Astra wins the raw-capability crown on some benchmarks, but Opus 5.5 costs less than half on every line — and on long cached-context agents, the input gap widens from 2.5x to roughly 20x per repeated token. Unless Astra measurably completes tasks Opus 5.5 fails, Opus 5.5 is the cheaper system.

Against its own sibling Claude Fable 5.1 ($10/$50): Fable is Anthropic's deepest-reasoning tier for work that must be exactly right; Opus 5.5 is the production agent tier. Same vendor, different jobs — route routine agentic traffic to Opus 5.5, reserve Fable for the hardest reasoning.

When Opus 5.5 is the right call

  • Long-running coding agents — stable repo context + repeated tool schemas = massive cache-hit rates
  • Customer-support agents with conversation memory — history re-sent every turn bills at $0.20/M
  • Document-grounded RAG systems — the same corpus chunks repeated across queries
  • Any workload where you already pay Opus 5 — it's a straight upgrade with a lower bill

When it's not

  • Single-shot bulk jobs with no repeated context — caching buys nothing, and cheaper models like DeepSeek V4.1 Flash or GPT-6 Luna win on sticker price
  • Latency-critical paths where a smaller model suffices

For a broader budget comparison, see the cheapest coding models ranked by cost per task, or the full live pricing dashboard.

Quick answers

How much does Claude Opus 5.5 cost?

Claude Opus 5.5 is listed at $4 per million input tokens and $20 per million output tokens as of September 2026, with prompt cache reads at $0.20 per million tokens. Released September 22, 2026.

Is Claude Opus 5.5 cheaper than Opus 5?

The list price is 20% below Opus 5 ($5/$25). Anthropic says typical long-context workloads cost around 40% less (Anthropic's measurement) thanks to lower token use and the cheaper $0.20/M cache reads.

How much do Claude Opus 5.5 cache reads cost?

Cache reads on Opus 5.5 cost $0.20 per million tokens — a 95% discount versus fresh input at $4. Agents with stable system prompts and conversation history routinely see cache-hit rates above 50%.

Is Opus 5.5 better than GPT-6 Astra for agentic work?

On price, decisively: Opus 5.5 at $4/$20 versus Astra at $10/$50. For long-running agents with cached context, Opus 5.5 is built specifically for that loop and often the cheaper pick per completed task.