Skip to main content
Same sticker price, a real-world cost drop from speed and token efficiency. Here is who should move today and who can wait a sprint.

Claude Sonnet 5.5: Should You Switch From Sonnet 5

KU
Kiril Urbonas
last week • 6 min read•1 view

Same sticker price, a real-world cost drop from speed and token efficiency. Here is who should move today and who can wait a sprint.

Key takeaways

  • Same sticker price, a real-world cost drop from speed and token efficiency.
  • Here is who should move today and who can wait a sprint.

Anthropic shipped Claude Sonnet 5.5 on September 28, the second model in the 5.5 line after Opus 5.5, and the pitch is unusual for a model release: the price per token did not move. Sonnet 5.5 still bills at $2 per million input tokens and $10 per million output, same as Sonnet 5, with cached input prefixes at $0.20 per million. The actual upgrade shows up on the invoice anyway, because the model runs over 30% faster and tends to finish the same task on noticeably fewer tokens, which Anthropic frames as up to 30% cheaper for typical token-billed workloads. If your Sonnet 5 traffic is agentic or terminal-heavy, this is worth migrating this week. If it is simple chat or extraction, you can wait a sprint and watch your own metrics first.

The benchmarks are real, but one number should make you suspicious#

Anthropic's own numbers: 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%. That is not a model getting a bit smarter, that is a nearly 7x jump, and jumps that size almost always mean the evaluation harness, scoring rubric, or tool scaffolding changed meaningfully alongside the model. Sonnet 5 was weirdly bad at Terminal-Bench 4.0 relative to how it performed in production, which suggests the earlier score was depressed by a harness or tooling mismatch as much as it reflects a genuine capability gap. Treat a 60-point swing as a flag to test your own tasks, not as a reason to trust the number at face value.

The rest of the benchmark sheet reads more like an ordinary point release. CursorBench 4.0 lands at 55.5%, about two points under Opus 5.5's 57.8%. OSWorld 2.1 computer-use scores 80.1%, close to Opus 5.5's 81.8%. Humanity's Last Exam with tools climbs to 64.5% from Sonnet 5's 54.9%, a meaningful but believable gain. Sonnet 5.5 is now sitting close enough to Opus 5.5 on agentic and computer-use work that the gap between the two tiers has mostly become a latency and cost decision, not a capability one.

Why the same price can mean a lower bill#

Token-metered billing means speed and verbosity are cost levers, not just UX levers. A model that reaches the right answer in fewer turns, with less restated context and fewer wasted tool calls, costs less per completed task even at an identical per-token rate. Sonnet 5 had a habit of re-reading large chunks of context or taking extra exploratory steps in long agent loops; if Sonnet 5.5 genuinely trims that, the saving shows up in your Anthropic invoice without you changing a single pricing tier.

bash.bash
# rough per-task cost check: compare Sonnet 5 vs 5.5 on the same prompt set
$ export ANTHROPIC_API_KEY=sk-ant-...
$ for MODEL in claude-sonnet-5 claude-sonnet-5-5; do
    echo "== $MODEL =="
    curl -s https://api.anthropic.com/v1/messages \
      -H "x-api-key: $ANTHROPIC_API_KEY" \
      -H "anthropic-version: 2026-01-01" \
      -H "content-type: application/json" \
      -d "{\"model\":\"$MODEL\",\"max_tokens\":1024,\"messages\":[{\"role\":\"user\",\"content\":\"$(cat task.txt)\"}]}" \
      | jq '{model: .model, input: .usage.input_tokens, output: .usage.output_tokens}'
  done

Run that against ten or twenty real prompts from your own traffic and sum the token counts before you believe any vendor's percentage. Anthropic's 30% figure is an average across their internal suite, and your workload is not their internal suite.

Who should move this week#

Anyone running agent loops that touch a terminal, a file system, or a browser was the worst-served audience on Sonnet 5, and Terminal-Bench's own 10.3% score backs that up even accounting for harness noise. If you have a CI agent, a coding assistant, or anything resembling the setups covered in running AI CLI agents in CI pipelines, swap the model identifier to claude-sonnet-5-5, rerun your eval set, and expect both a quality and a cost improvement in the same move. Computer-use and long multi-step workflows are the other clear win, since the gap to Opus 5.5 on OSWorld 2.1 has nearly closed; teams paying for Opus just to get reliable computer-use behavior should re-test on Sonnet 5.5 before renewing that spend.

Who can wait#

If your Sonnet 5 usage is mostly single-turn chat, summarization, or structured extraction, the delta is smaller and less certain. Those workloads were never where Sonnet 5 struggled, so you are less likely to see the dramatic token savings that show up in long agent loops. Run the comparison above on a sample of your real traffic before flipping the default model in production. A second reason to wait: new model identifiers occasionally produce subtly different formatting or tool-call conventions, and a wholesale swap without a canary period is how a working pipeline quietly starts failing.

This is not a new tier, it's the same tier working better#

Nothing about Sonnet 5.5 changes where it sits in Anthropic's lineup. It is still the mid-tier model between Haiku and Opus, still priced at $2/$10, and still the default choice for most product traffic per our broader take in best LLM APIs by use case. What changed is how much work that tier does per dollar, which is exactly the kind of improvement that matters more than a benchmark headline: it does not require you to re-architect your routing logic in OpenAI vs Anthropic vs Gemini, it just makes the Anthropic leg of that routing cheaper to run.

The decision, concretely#

  • Does your Sonnet 5 traffic run agentic or terminal-heavy tasks? Switch to Sonnet 5.5 now and re-run your eval set; this is the clearest win in the release.
  • Are you paying for Opus 5.5 mainly for computer-use reliability? Re-test on Sonnet 5.5 first; the OSWorld 2.1 gap (80.1% vs 81.8%) is close enough that you may not need the premium tier anymore.
  • Is your traffic mostly single-turn chat or extraction? Sample your own prompts through the cost check above before flipping the production default; the token savings there are smaller and unverified for your workload.
  • Does a benchmark jump this large (10.3% to 70.6%) make you want to trust the headline number alone? Don't. Validate against your own tasks before treating it as a measured capability gain.

The call we'd make#

Move agentic, coding, and computer-use workloads to Sonnet 5.5 this week; the token-efficiency gains and the near-Opus computer-use score make it a low-risk swap at an unchanged price. For chat and extraction traffic, run a one-week canary against your own prompts before making it the default, because the headline savings are an average, not a guarantee for every workload. Either way, re-test rather than assume: a benchmark swing this large deserves your own eval before it touches production traffic.

Explore topics:AI
React

Get the DevOps Troubleshooting Cheat Sheet

Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.

Share this post
KU

Kiril Urbonas

AI Engineer

560 articles
View all articles by Kiril Urbonas

You might have missed

Evergreen posts worth revisiting.