Claude Opus 5 and **📝Claude Opus 5.5** are consecutive Opus-tier models from 📝Anthropic, released two months apart with the same 1M-token context window and 128K output limit. They diverge in how they think and how they are controlled: Opus 5.5 makes thinking always on, lowers the default effort to medium, costs 20% less per token, and adds Mythos-class safeguards, which makes it a different model to operate rather than a drop-in upgrade.
At a Glance
Thinking and Effort
Claude Opus 5 accepts thinking: {"type": "disabled"} at effort high or below, so latency-sensitive integrations could switch reasoning off. Claude Opus 5.5 rejects that setting and any manual budget_tokens with a 400 error; effort is the only control. Effort levels are not calibrated the same way across the two models. According to Anthropic's prompting guide, Opus 5.5 at medium matches or exceeds Opus 5 at high on coding and knowledge-work evaluations, yet at any given level Opus 5.5 thinks more per turn, especially at xhigh and max. An integration that carries its Opus 5 effort setting forward will see longer turns and higher output-token counts, so the effort sweep has to be re-run rather than copied.
Cost and Speed
Per-token prices fall 20%, from $5 / $25 to $4 / $20 per million tokens, and cached reads fall from $0.50 to $0.20. Anthropic reports that Opus 5.5 costs 40% less than Opus 5 to run on typical workloads, because it also generates output tokens more than 30% faster and tends to finish the same task with fewer tokens. The saving depends on effort choice: running Opus 5.5 at its default medium rather than Opus 5's default high is where most of it comes from.
Agent Behavior and API Surface
Opus 5.5 changes how agent harnesses must read responses. Notes written between tool calls arrive as thinking blocks rather than text blocks, empty unless thinking.display is set to "updates", so clients built for Opus 5 can appear silent. Opus 5.5 also ends some turns with a progress report instead of a tool call, which an unattended loop can mistake for completion. On the 📝Claude API and 📝Google Cloud, computer use requires the computer_toolset_20260801 toolset, and the older computer_20251124 tool is rejected. Forced tool_choice values of any or tool return errors, with strict tool use as the replacement.
Reported Benchmarks
Anthropic's launch figures for Opus 5.5 versus Opus 5: Terminal-Bench 4.0, 66.4% versus 52.3%; CursorBench 4.0, 57.8% versus 46.6%; FrontierCode v1.1, 54.4% versus 48.0%; OSWorld 2.1, 81.8% versus 74.0%; and GDPval-AA v2.1, 1846 versus 1708 Elo. These are vendor-reported and were not independently reproduced for this memo.
Where Each Wins
Claude Opus 5 wins when an integration depends on behavior Opus 5.5 removed: thinking disabled for the lowest possible latency, forced tool_choice, or the computer_20251124 tool on the Claude API without time to migrate. It also remains the right choice for life-sciences workloads that Opus 5.5's new biology classifier blocks, and it serves as a permitted fallback model when Opus 5.5 or Claude Fable 5.1 decline a request.
Claude Opus 5.5 wins for nearly everything else: long-running agentic coding and code review, knowledge work where factual accuracy matters, dense charts and screenshots, and computer use. It is cheaper per token, faster per task, and, according to Anthropic, close to Claude Fable 5.1 on most work.
Related
- 📝Claude Opus 5.5 — reference memo with specs and FAQ
- 📝How to Prompt Claude Opus 5.5 — symptom-first fixes for migrating prompts
- 📝Claude Model Glossary — every Claude model by tier
- 📝Claude Mythos — the tier whose safeguards Opus 5.5 inherits
- 📝Context Engineering Rules for Claude 5 Generation — context practices for both models
We run BotBrian on Opus 5.5, and the lesson from reading both models' documentation side by side is that the version number undersells the change. Re-run effort settings, check how your client renders thinking blocks, and treat anything that relied on switching thinking off as needing redesign.
