Sonnet 5 had a rough launch over token usage and a forced default, not weak capability. One important fact changed after publication: Anthropic cancelled the planned price increase. Since August 10, Sonnet 5 has stayed at $2 per million input tokens and $10 per million output tokens. The lower rate helps, but long agent runs can still cost more than expected because this model may emit more tokens and take more steps. The benchmark and cost-per-task sections below describe the launch-period evidence; the price and product details are current as of September 1.
- Anthropic released Claude Sonnet 5 on June 30, 2026, and made it the default model across plans.
- Anthropic made $2 per million input tokens and $10 per million output tokens the standard Sonnet 5 API price on August 10, 2026.
- Anthropic calls it its most agentic Sonnet yet. TechCrunch framed the launch as a cheaper way to run agents.
- On Anthropic's own benchmarks it beats Sonnet 4.6 across the board and edges Opus 4.8 on the GDPval-AA v2 knowledge-work test, 1,618 to 1,615.
- Independent testing by Artificial Analysis corroborates the capability but measured about $2.29 per task, roughly 15% more than Opus 4.8, because Sonnet 5 uses far more tokens.
- A new tokenizer emits about 30% more tokens for the same text, per Simon Willison, so the lower sticker price does not translate one to one into a lower bill.
- The launch-week backlash, June 30 to July 2, was mostly about cost and the forced default, not capability. Hands-on coding reviews a week later are largely positive.
The short version of the Sonnet 5 story is that the anger was real and the reason was misread. People saw a wave of complaints and assumed the model was bad at its job. It is not. What set people off was the token bill and a new default they did not choose. Those are fair complaints. They are also different from "this model cannot code." Here is the honest version.
What changed in Sonnet 5
The most agentic Sonnet yet, and now the default.
Anthropic released Claude Sonnet 5 on June 30, 2026 and describes it as its most agentic Sonnet, tuned for reasoning, tool use, and long multi-step coding tasks. It became the default across plans on the same day, replacing Sonnet 4.6 for many users. That switch explains the launch reaction. It should not be used as a current model-picker reference: consumer and API availability change separately, so check Anthropic's current model overview before choosing an older model for a workflow.
The price story changed after launch. Anthropic first called $2 per million input tokens and $10 per million output tokens an introductory rate. On August 10, it made that rate standard instead of raising Sonnet 5 to $3 and $15. That correction matters: every estimate in this review should now start from $2 and $10. Long runs still need measurement because the number of billed tokens can erase part of the rate advantage.
Two quieter changes matter more than the price tag. Sonnet 5 ships a new tokenizer, and its adaptive thinking is more verbose and takes more steps. Both of those turn into tokens, and tokens are what you pay for. Hold that thought, because it is the whole cost story.
The benchmarks Anthropic led with
Strong numbers, with an asterisk on how they were chosen.
On Anthropic's own evaluations, Sonnet 5 is a real step up from Sonnet 4.6 and comes surprisingly close to Opus 4.8. The most talked-about result is knowledge work, where Sonnet 5 edges the concurrent Opus flagship. This is reportedly the first time a mid-tier Sonnet has beaten the Opus released alongside it on any benchmark.
Claude Sonnet 5 vs Sonnet 4.6 vs Opus 4.8, Anthropic's figures
| Sonnet 4.6 | Sonnet 5 | Opus 4.8 | |
|---|---|---|---|
| SWE-bench Pro, agentic coding | 58.1% | 63.2% | 69.2% |
| OSWorld-Verified, computer use | 78.5% | 81.2% | not shown |
| GDPval-AA v2, knowledge work (Elo) | not shown | 1,618 | 1,615 |
Source: Anthropic, Claude Sonnet 5 announcement. Figures are vendor-reported. Not shown means the number was not part of Anthropic's own comparison.
Read the coding row carefully, because it is the honest frame for the whole model. Sonnet 5 clearly beats Sonnet 4.6, and it clearly trails Opus 4.8 on pure coding. It is not a new frontier leader. It is a cheaper model that got close.
One asterisk on the coding claim
Anthropic led with SWE-bench Pro, the harder and newer variant, where Sonnet 5 comes out ahead of GPT-5.5. On the more common SWE-bench Verified, third-party writeups place GPT-5.5 and Gemini 3.1 Pro higher. So "best coder" depends on which leaderboard you pick. The independent eval site Vellum also notes Anthropic restated some Sonnet 4.6 scores after a methodology change, which is worth knowing before you treat these as clean apples-to-apples numbers.
The cost catch nobody mentioned at launch
A lower sticker price is not a lower bill.
This is the part that earned the backlash, and it is the most useful section of this review. The pitch is that Sonnet 5 is the cheap way to run agents. The reality, once independent testers measured it, is that a lower per-token rate does not mean a lower total.
Start with the tokenizer. Simon Willison measured the new tokenizer producing about 30% more tokens for the same text, roughly 1.4 times for English, 1.33 times for Spanish, and 1.27 times for Python, with Mandarin about unchanged. Every one of those extra tokens is billed. So the same task quietly costs more than the rate card suggests, before the model has done anything different.
Then add behavior. Sonnet 5 thinks more and takes more steps. Artificial Analysis, which had independent pre-release access, found Sonnet 5 uses roughly 40% more output tokens per task, and put a hard number on the result: about $2.29 per task, which is around double Sonnet 4.6 and roughly 15% more than Opus 4.8. The model designed to undercut Opus measured more expensive than Opus on their agentic workload. Vellum reached the same conclusion in different words, noting that at high effort settings Sonnet 5 can cost more than Opus 4.8 for similar quality.
Not sure which AI model to use?
21 models · Personalized picks · 60 seconds
When the discount is real, and when it is not
Sonnet 5's $2 / $10 rate is now permanent. It is usually easier to justify on short, low-effort tasks. On long, high-effort agent loops, extra tokens and turns can still push the total above a model with a higher sticker price. If you run autonomous jobs, price the workload instead of comparing rate cards alone.
If you want to sanity-check your own numbers, our AI cost calculator lets you compare per-model spend, and the benchmark leaderboard tracks how the current models score with the source behind each cell.
So where does Sonnet 5 rank?
A strong tier-two model, not a new frontier leader.
Put the launch noise aside and the independent picture is consistent. Artificial Analysis placed Sonnet 5 at an Intelligence Index of 53, around fifth overall, only a couple of points behind GPT-5.5 at high effort and Opus 4.8 at max effort, and about six points above Sonnet 4.6. It leads its price class on agentic coding and it genuinely edges Opus 4.8 on the GDPval knowledge-work measure, which matches Anthropic's own claim.
What it is not is a generational leap. A common line among frontier watchers was that this feels like a Sonnet 4.8 or 4.9, not a 5.0. That perception was not helped by timing, since Anthropic launched Fable 5 the very next day and pulled attention straight to it. If you want the wider context on that release, see our writeup on the Fable 5 and Mythos 5 situation.
The other gripes that stuck
Overthinking, and an upgrade nobody asked for.
Two more complaints show up often enough to be real, not noise. The first is that Sonnet 5 overthinks small tasks. Where Sonnet 4.6 hands back a quick answer, Sonnet 5 keeps working, which is great for a refactor and annoying for a one-line fix. AlphaSignal put it as a recommendation: use Sonnet 5 for serious multi-file work, keep 4.6 for quick edits. CodeRabbit, which is otherwise very positive, measured a drop in bug-detection recall versus 4.6 even as precision went up, and the sharpest launch-week review, Every's Vibe Check, said it can get stuck in loops on complex coding. Balance that against the praise, because the same reviewers who flag the overthinking also call it a clear step up for real building work.
The second is the forced default. Sonnet 5 replaced Sonnet 4.6 across the consumer plans on June 30, and 4.6 was pulled from the model picker on claude.ai, so users who preferred its quick and cheaper behavior could not simply keep selecting it. Sonnet 4.6 still runs through the API, but for many Free and Pro users the everyday model changed under them without a say. A lot of the "this feels worse" energy in the first days came from people comparing an unfamiliar default to a model they had tuned their habits around. Anthropic also deliberately limited Sonnet 5 on some math and security tasks, and its own system card shows a small uptick in refusals versus 4.6, low in absolute terms but concentrated in sensitive domains.
Who should use it
It depends on the job and the length of the run.
Match the model to the task
- 1Everyday agentic coding, refactors, and browser research: Sonnet 5 is a strong default at its permanent $2 / $10 API rate.
- 2Long, high-effort autonomous agent loops: watch the token bill. Independent tests put Sonnet 5 above Opus 4.8 on cost per task once runs get long.
- 3Quick one-off edits: Sonnet 4.6 was often faster in launch-period reviews. Check Anthropic's current model overview before assuming it is available on your plan or API account.
- 4Deep math or security research: test Sonnet 5 against Anthropic's current Opus tier on your own workload instead of relying on this launch-period Opus 4.8 comparison.
If you are weighing Sonnet 5 against the wider field rather than just its Anthropic siblings, our task-by-task model guide covers where the flagships land on coding, agents, and cost, and if the pull is toward cheaper coding specifically, the best open-source coding models guide and our Claude Code vs Codex breakdown are the natural next reads.
The verdict
Overblown reputation, earned cost skepticism.
Claude Sonnet 5 did not get bad reviews for being a bad model. It had a bad launch week over price and a forced default, and the capability reviews that followed are positive. It beats Sonnet 4.6 for real building, it flirts with Opus 4.8 on agent benchmarks, and it edges Opus on knowledge work. The one criticism that survives contact with the evidence is the cost story, and it is a good one: the model sold as the cheaper way to run agents can cost more than the model it was meant to undercut, once the tokenizer and the extra turns are counted.
The honest summary: Sonnet 5 is easier to recommend now that $2 and $10 are permanent. Use it for everyday agentic coding, compare it with Anthropic's current Opus tier for long or high-stakes work, and measure token use before moving a large workload. The launch-week evidence still explains its reputation; it no longer explains its current price.
Sources
- Anthropic: Introducing Claude Sonnet 5
- Anthropic API release notes: Sonnet 5 standard pricing update
- Anthropic: What's new in Claude Sonnet 5
- Anthropic: Sonnet 5 migration guide
- Simon Willison: Claude Sonnet 5, and the new tokenizer
- Artificial Analysis: Claude Sonnet 5 model page and cost-per-task testing
- CodeRabbit: Claude Sonnet 5 hands-on code-review evaluation
- Vellum: Claude Sonnet 5 benchmarks explained
- TechCrunch: Anthropic launches Claude Sonnet 5 as a cheaper way to run agents
FAQ
Did Claude Sonnet 5 really get bad reviews?
Only for its first two days, and mostly about cost rather than capability. The June 30 to July 2 backlash focused on a new tokenizer that inflates token counts, a cost per task that independent testing put above Opus 4.8, and Sonnet 5 being made the forced default on Free and Pro plans. Hands-on coding reviews that landed about a week later were largely positive, with reviewers like CodeRabbit calling it the most exciting coding model in its class.
Is Claude Sonnet 5 cheaper than Opus 4.8?
On the sticker price, yes: Anthropic made $2 per million input tokens and $10 per million output tokens the standard rate on August 10. The actual bill still depends on token use and agent turns. Independent launch-period testing measured about $2.29 per task, roughly 15% more than Opus 4.8 on that evaluator's workload.
Is Sonnet 5 better than Sonnet 4.6?
For real building work, most reviewers say yes. Sonnet 5 beats 4.6 across Anthropic's own benchmarks and reviewers describe it as a clear step up for multi-file and agentic coding. The trade-off is that it tends to overthink small tasks, so for quick one-off edits Sonnet 4.6 is often faster and cheaper. CodeRabbit also measured a drop in bug-detection recall versus 4.6 even as precision improved.
Why does Claude Sonnet 5 use more tokens?
Two reasons. First, Sonnet 5 ships a new tokenizer that, by Simon Willison's measurements, produces about 30% more tokens for the same text, roughly 1.4 times for English and 1.27 times for Python. Second, its adaptive thinking is more verbose and it takes more steps on agent tasks. Together these can offset or exceed the lower per-token price on long runs.
Can I switch back to Sonnet 4.6?
Do not rely on an old claude.ai model-picker list. Anthropic changes consumer availability separately from API availability. Check the current picker for your plan and Anthropic's official model overview before building a workflow around Sonnet 4.6. The June 30 switch still explains why the launch felt abrupt; it does not prove today's lineup.

Founder of Spectrum AI Labs — testing AI tools and models, and writing up what actually ships.
More about Paras →Stay ahead of the AI curve
We test new AI tools every week and share honest results. Join our newsletter.



