The headline numbers
Anthropic's listed prices for Claude Opus 5.5, verified September 26, 2026:
| Token type | Price / 1M | What it is |
|---|---|---|
| Input (fresh) | $4.00 | New tokens never seen before |
| Input (cache read) | $0.20 | Repeated context — 95% discount |
| Output | $20.00 | Generated tokens, incl. thinking |
That $0.20 cache-read line is the whole story of this model. Everything else — the $4/$20 headline, the release timing, the agent positioning — is context for why cache reads matter so much here.
20% off the sticker, ~40% off typical bills
Opus 5.5's list price sits 20% below Opus 5 ($5/$25). Anthropic says the combination of lower rates and lower token use reduces typical workload cost by around 40% versus Opus 5 — a vendor measurement, not a guarantee for every workload. Two factors drive it beyond the raw discount:
- Lower token use in real testing. GitHub's chief product officer said Opus 5.5 used among the fewest tokens and steps in GitHub's Copilot CLI and VS Code testing — solving more terminal tasks than Opus 5 while using less than half as many steps — so the same task costs fewer billable tokens.
- Cheaper cache reads compounding. At $0.20/M — down from Opus 5's higher cache rate — every repeated prompt segment costs almost nothing. In workloads with stable system prompts, the cached portion of the bill collapses.
This is exactly the "token price isn't task price" effect: a 20% sticker cut becomes a 40% bill cut because the cheaper model also consumes fewer billable tokens per task.
Cache-read economics: where agents win
Prompt caching works like this: when you resend context the model has already seen — system prompts, tool definitions, conversation history — the provider skips reprocessing it and bills those tokens at the cache-read rate instead of the fresh-input rate. On Opus 5.5 that means $0.20/M instead of $4.00/M: a 95% discount.
Now consider a coding agent's actual loop. Each step of the loop re-sends the full conversation: the repo context, the system prompt, the tool schemas. In production agent deployments, cache-hit rates above 50% are common — often 70–90% once the context stabilizes.
A worked example
A coding agent session burns 2M input tokens total, with 80% cache hits, and generates 200K output tokens:
Cached input: 1.6M × $0.20 = $0.32
Output: 0.2M × $20.00 = $4.00
Total: $5.92 for the session
Without caching, the input line alone would be $8.00 instead of $1.92. The cache turns a $12 session into a $5.92 session. This is why Opus 5.5 is positioned as the agent-economics pick: the longer the agent runs, the larger the cached share, and the more the effective price approaches the cache-read rate rather than the headline rate.
Opus 5.5 vs the field
| Model | Input / 1M | Output / 1M | Cache read / 1M |
|---|---|---|---|
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 |
| GPT-6 Astra | $10.00 | $50.00 | discounted* |
| GPT-6 Sol | $2.00 | $10.00 | discounted* |
| Claude Fable 5.1 | $10.00 | $50.00 | $0.25 |
| Grok 4.7 | $2.00 | $6.00 | $0.50 |
*OpenAI discounts cached input on the GPT-6 family; check current docs for exact rates.
The comparison that matters: Opus 5.5 vs GPT-6 Astra at $10/$50. Astra wins the raw-capability crown on some benchmarks, but Opus 5.5 costs less than half on every line — and on long cached-context agents, the input gap widens from 2.5x to roughly 20x per repeated token. Unless Astra measurably completes tasks Opus 5.5 fails, Opus 5.5 is the cheaper system.
Against its own sibling Claude Fable 5.1 ($10/$50): Fable is Anthropic's deepest-reasoning tier for work that must be exactly right; Opus 5.5 is the production agent tier. Same vendor, different jobs — route routine agentic traffic to Opus 5.5, reserve Fable for the hardest reasoning.
When Opus 5.5 is the right call
- Long-running coding agents — stable repo context + repeated tool schemas = massive cache-hit rates
- Customer-support agents with conversation memory — history re-sent every turn bills at $0.20/M
- Document-grounded RAG systems — the same corpus chunks repeated across queries
- Any workload where you already pay Opus 5 — it's a straight upgrade with a lower bill
When it's not
- Single-shot bulk jobs with no repeated context — caching buys nothing, and cheaper models like DeepSeek V4.1 Flash or GPT-6 Luna win on sticker price
- Latency-critical paths where a smaller model suffices
For a broader budget comparison, see the cheapest coding models ranked by cost per task, or the full live pricing dashboard.
Quick answers
How much does Claude Opus 5.5 cost?
Claude Opus 5.5 is listed at $4 per million input tokens and $20 per million output tokens as of September 2026, with prompt cache reads at $0.20 per million tokens. Released September 22, 2026.
Is Claude Opus 5.5 cheaper than Opus 5?
The list price is 20% below Opus 5 ($5/$25). Anthropic says typical long-context workloads cost around 40% less (Anthropic's measurement) thanks to lower token use and the cheaper $0.20/M cache reads.
How much do Claude Opus 5.5 cache reads cost?
Cache reads on Opus 5.5 cost $0.20 per million tokens — a 95% discount versus fresh input at $4. Agents with stable system prompts and conversation history routinely see cache-hit rates above 50%.
Is Opus 5.5 better than GPT-6 Astra for agentic work?
On price, decisively: Opus 5.5 at $4/$20 versus Astra at $10/$50. For long-running agents with cached context, Opus 5.5 is built specifically for that loop and often the cheaper pick per completed task.