Claude Sonnet 5.5 vs Opus 5.5: Which One Fits Your Work?

One week, two models — and the cheaper one is uncomfortably close to the flagship. On September 22, 2026, Anthropic released Claude Opus 5.5, the first model in its new Claude 5.5 family and, in the company's words, its most powerful model to date. Exactly one week later, the second family member arrived: Claude Sonnet 5.5, aimed at the well-scoped everyday work that fills most of the workday. The sticker prices are half a tier apart — $4 versus $2 per million input tokens. But according to Anthropic's own numbers, Sonnet 5.5 now matches or beats Opus 5.5 on several of the benchmarks developers care about. Which one should you actually use?
What just happened
Claude Opus 5.5 launched on September 22, 2026. Anthropic describes it as its most powerful AI model to date, performing on par with its top-tier Claude Fable 5.1 on most tasks while costing 40% less to run than Opus 5 on typical workloads, the company claims. API pricing is $4 per million input tokens and $20 per million output tokens, with cheaper caching ($0.20 per million cache reads, $5 per million cache writes). It is available on AWS, Google Cloud, Microsoft Azure, and the Claude Platform as claude-opus-5-5, with zero data retention.
Claude Sonnet 5.5 launched on September 28, 2026, as the second model in the Claude 5.5 family. Anthropic positions it as a faster, cheaper complement to Opus 5.5, strongest at everyday tasks with clear scope: fixing bugs, creating polished documents, slides, and spreadsheets. The model keeps Sonnet 5's pricing — $2 per million input tokens, $10 per million output tokens, $0.20 per million cache reads — but, according to Anthropic, typically needs far fewer tokens to complete the same work, which the company says cuts costs by up to 30% per task. Output generation is reportedly more than 30% faster than Sonnet 5, making it the fastest Sonnet model to date. It is available across all major cloud platforms as claude-sonnet-5-5, and a third model, Haiku 5.5 for fast, high-volume work, is expected soon.
The price gap, side by side
| Claude Opus 5.5 | Claude Sonnet 5.5 | |
|---|---|---|
| Input price (per 1M tokens) | $4 | $2 |
| Output price (per 1M tokens) | $20 | $10 |
| Cache reads (per 1M) | $0.20 | $0.20 |
| Cache writes (per 1M) | $5 | $2.50 |
| Speed claim | 30%+ faster output than Opus 5 | 30%+ faster output than Sonnet 5 |
| Efficiency claim | 40% cheaper to run than Opus 5 | Up to 30% less cost per task than Sonnet 5 |
| Launch date | September 22, 2026 | September 28, 2026 |
All figures and efficiency claims come from Anthropic and have not yet been independently verified. One subtlety: Sonnet 5.5's "up to 30% cheaper per task" is not a price cut — it is an efficiency claim based on the model using fewer tokens and fewer tool-call steps than its predecessor. For developers running agentic workloads, where token volume is the real bill, that distinction matters more than the sticker price.
The benchmarks — with the vendor asterisk
Here is where the comparison gets interesting. All scores below are company-reported; treat them as vendor numbers until independent evaluations land.
On agentic coding tests, Sonnet 5.5 holds its own against the flagship tier. On Terminal-Bench 4.0, Sonnet 5.5 scores 70.6% versus Opus 5.5's 66.4% (and Sonnet 5's 10.3%) — though Anthropic's own footnote notes Opus 5.5's score is its best at the "Xhigh" effort setting. On CursorBench 4.0, which recreates real coding sessions from the Cursor editor, Sonnet 5.5 scores 55.5%, just two points below Opus 5.5's 57.8%, and far above Sonnet 5's 34.1%. On FrontierCode 1.1, Sonnet 5.5 reaches 52.1% at the "Xhigh" setting against Opus 5.5's 54.4% — but drops to 46.2% at maximum effort, because, Anthropic says, at "Max" the model more often triggers a code-review function that causes timeouts the benchmark penalizes.
On knowledge work, the gap is nearly gone. On GDPval-AA, an OpenAI-developed benchmark covering 44 professions, Sonnet 5.5 scores 1,844 points against Opus 5.5's 1,846 — well ahead of Sonnet 5's 1,449. Anthropic flags that these numbers came from a pre-release build in which a since-fixed bug may have affected structured outputs. On Chartography, a visual chart-recognition test, Sonnet 5.5 jumps to 61.6% from Sonnet 5's 15.6%, close to Opus 5.5's 64.4%. On Humanity's Last Exam with tools, Sonnet 5.5 scores 64.5% versus Opus 5.5's 67.7%; on OSWorld 2.1 computer skills, 80.1% versus 81.8%.
The honest read: on the everyday tasks these models will actually do — agentic coding, knowledge work, document polish — the difference is a few percentage points on vendor benchmarks. Opus keeps a small edge in interdisciplinary thinking and computer use. With margins this tight, per-task cost may be the deciding factor.
Who should use which
Choose Opus 5.5 when the work demands careful judgment: complex multi-step projects, long-running agentic tasks, and situations where a wrong answer is expensive. Anthropic built it for complex work requiring careful judgment, and it is the company's most capable model for enterprise deployments, with external evaluations by METR and Frontier Design before launch and safeguards previously reserved for its most powerful systems. If your workload already pushes Sonnet's limits, Opus 5.5 is the escalated tier.
Choose Sonnet 5.5 for everyday, well-scoped work: fixing bugs, reviewing pull requests, drafting and polishing documents, slides, and spreadsheets. According to Anthropic, Sonnet 5.5 batches tool calls more efficiently than its predecessor, cutting the number of steps per task — which is why the company says costs drop by up to 30% per task even though the sticker price is unchanged. Early testers praised its feel for design, the company claims: reworking user interfaces and executing slide templates so well that results barely need touch-ups. The 30%-plus speed gain makes it the better fit for fast iteration loops.
A workable rule of thumb: default to Sonnet 5.5, and escalate to Opus 5.5 only for the tasks where judgment quality — not speed or cost — is the bottleneck. And if you run AI agents at volume, benchmark your own workloads: measure cost per completed task, not cost per token. That is the metric both tiers are now competing on.
Safety and safeguards
The launches arrive at an unusual moment. Weeks earlier, Anthropic CEO Dario Amodei called on the AI industry to slow down frontier releases so safety practices could keep up — a call echoed by OpenAI's Sam Altman and Elon Musk. Anthropic was explicit that Sonnet 5.5 "doesn't advance the frontier of our model's capabilities," which is why no new guardrails were added. At the same time, Sonnet 5.5 is the first Sonnet model to launch with cybersecurity safeguards previously used only for Anthropic's most capable models: requests involving high-risk cybersecurity tasks get visibly rerouted to Sonnet 5, while vetted experts can apply for tiered access through an expanded Cyber Verification Program. The model also adds classifiers against distillation attacks — the same protections used in Anthropic's top models — which the company ties to EU AI Act rules on systemic risk.
For Opus 5.5, the safety picture was covered in our earlier comparison: external evaluation by independent groups, zero data retention, and candid admissions in the system card about regressions — the model is more likely than previous versions to follow malicious instructions pasted into a prompt. The practical takeaway for enterprise buyers is that safeguards are improving but failure modes are shifting, not disappearing, and organizations with sensitive use cases should keep the verification programs on their radar.
How to try them
- API access. Select
claude-sonnet-5-5orclaude-opus-5-5on the Claude Platform, AWS Bedrock, Google Cloud Vertex AI, or Microsoft Azure. Enable prompt caching — cache reads are $0.20 per million tokens on both models, and Anthropic says caching accounts for the majority of agentic and coding-work costs. - Start with Sonnet 5.5 for everyday work and coding tasks. Its combination of speed, lower per-task cost, and near-flagship benchmark scores makes it the sensible default.
- Tune the effort setting. Like Anthropic's other recent models, Sonnet 5.5 offers an adjustable effort setting to trade cost and speed against quality. Anthropic says that at low or medium effort, Sonnet 5.5 already beats Sonnet 5's best scores on several benchmarks at about one-tenth the per-task cost — but avoid "Max" for agentic coding until the sub-agent timeout wrinkle is resolved.
- Subscription users. Check your plan's updated usage limits after the Opus 5.5 launch; Anthropic has been raising five-hour limits on Pro, Max, Team, and Enterprise plans.
- Explore the landscape. If you are deciding between assistant ecosystems rather than API tiers, browse our AI chatbots directory for the tools that run on these models.
The bottom line
September was less a capability leap than a tier-structure shift: the honest gap between Anthropic's "flagship" and "workhorse" models has shrunk to a few percentage points on the company's own benchmarks, while the price gap between them is 2x per token. For most everyday coding and knowledge work, Sonnet 5.5 is now the sensible default — faster, cheaper per task, and within rounding error of Opus 5.5 where it counts. Opus 5.5 remains the right call when judgment quality is the bottleneck and a wrong answer is expensive.
Keep the vendor claims in perspective, though. "Up to 30% cheaper per task," "matches Opus within two points," "fastest Sonnet to date" — all company-reported, all awaiting independent confirmation, and at least one flagged with a pre-release caveat. The direction is clear; the exact numbers deserve verification. Test both against your own tasks, and let cost per completed task — not benchmark bragging rights — decide.
Sources
- Reuters, "Anthropic rolls out second Claude 5.5 model as it builds toward IPO," September 28, 2026. https://www.reuters.com/technology/anthropic-rolls-out-second-claude-55-model-it-builds-toward-ipo-2026-09-28/
- THE DECODER, "Anthropic's Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks while costing up to 30 percent less per task," September 28, 2026. https://the-decoder.com/anthropics-claude-sonnet-5-5-nearly-matches-opus-5-5-on-benchmarks-while-costing-up-to-30-percent-less-per-task/
- AWS News Blog, "AWS Weekly Roundup: GPT-6 Sol and Luna, Claude Opus 5.5 on Amazon Bedrock, Strands harness, and more (September 28, 2026)," September 28, 2026. https://aws.amazon.com/blogs/aws/aws-weekly-roundup-gpt-6-sol-and-luna-claude-opus-5-5-on-amazon-bedrock-strands-harness-and-more-september-28-2026/