Skip to main content
Haiku 5.5 is cheap enough to stop rationing subagent calls, but the 75% savings claim only holds if your prompts are short.

Claude Haiku 5.5: When to Route Down From the Frontier Tier

KU
Kiril Urbonas
4 days ago • 6 min read•0 views

Haiku 5.5 is cheap enough to stop rationing subagent calls, but the 75% savings claim only holds if your prompts are short.

Key takeaways

Haiku 5.5 is cheap enough to stop rationing subagent calls, but the 75% savings claim only holds if your prompts are short.

Anthropic shipped Claude Haiku 5.5 on October 7, 2026, priced at $0.10 per million input tokens and $0.50 per million output for prompts up to 100K tokens, rising to $0.50 and $2.50 above that, versus Haiku 4.5's flat $1/$5. The headline is a claimed 75% drop in average cost to run, and the real question isn't whether Haiku 5.5 is cheap, it's which of your current Sonnet or GPT calls never needed the frontier tier in the first place. We think that's most of your subagent fan-out, almost none of your user-facing judgment calls, and the honest way to find the line is to measure your own workload instead of trusting the vendor's blended number.

The 75% number is a tokenizer trick, not a free lunch#

Haiku 5.5 reportedly ships a new tokenizer that counts the same text as noticeably more tokens than Haiku 4.5's did, with some reports putting the inflation around 30%. That means a prompt sitting comfortably under the old model's token count can cross the 100K threshold on the new one and get billed at the higher tier, $0.50/$2.50 instead of $0.10/$0.50. Anthropic's 75% figure is reportedly a blended average weighted toward the real-world mix of short requests, which is a fair way to describe a fleet, but it's not a promise about any single call. Before you re-price a pipeline, run your actual prompts through the new tokenizer and look at where they land relative to 100K. A summarization job that was 70K tokens on Haiku 4.5 might be 90K now, still fine. A RAG-stuffed context that was 95K might tip over into the expensive tier and erase most of the savings.

Where routing down is an easy call#

Subagent fan-out is the clearest case. If your orchestrator spins up ten parallel calls to classify tickets, extract fields, or decide whether a log line matters, each one is cheap to get wrong and cheap to retry. Haiku 5.5's new effort controls reportedly let you dial capability against token spend per request, so a trivial classification can run at low effort while a messier one escalates, without hand-rolling that logic yourself. Browser-use loops are the same shape: lots of small, recoverable decisions (click this, read that, retry) where a 20% miss rate just means a few more steps, not a wrong answer reaching a customer.

python.python
# router.py - pick a model tier by task class, not by habit
from dataclasses import dataclass

@dataclass
class TaskSpec:
    task_class: str          # "subagent", "summarize", "extract", "judge", "codegen"
    stakes: str              # "low", "medium", "high"
    prompt_tokens: int
    retry_is_cheap: bool

CHEAP_MODEL = "claude-haiku-5-5"
FRONTIER_MODEL = "claude-sonnet-5-5"

def choose_model(spec: TaskSpec) -> str:
    # high-stakes or expensive-to-retry work never routes down
    if spec.stakes == "high" or not spec.retry_is_cheap:
        return FRONTIER_MODEL

    # long prompts erase Haiku 5.5's pricing edge past the 100K tier
    if spec.prompt_tokens > 90_000:
        return FRONTIER_MODEL

    if spec.task_class in {"subagent", "summarize", "extract"}:
        return CHEAP_MODEL

    # judge and codegen tasks default to frontier unless proven otherwise
    return FRONTIER_MODEL

That function is deliberately boring. The interesting part is retry_is_cheap: it's the actual variable that should decide tier, not task labels, because the same task class can be low-stakes in one product and high-stakes in another.

Where it isn't a close call at all#

Anything where a wrong answer costs money, time, or trust stays on the frontier tier. Code generation that gets merged without a human diff review, financial or legal extraction, final-answer customer support, multi-step agent planning where an early mistake compounds through every later step: none of that is a 20%-miss-is-fine situation. We've also found Haiku-class models degrade faster than Sonnet-class ones on long, ambiguous instructions, so a task that's cheap today can get expensive the moment the prompt grows a few more edge cases. If you're still deciding between the two frontier options for that work, our take on switching to Claude Sonnet 5.5 and the GPT-6.1 Sol pricing cut both cover that comparison directly. This post isn't trying to repeat either one. It's the layer underneath: deciding whether you need a frontier model at all before you decide which one.

Effort controls change the routing math, carefully#

The effort slider is the more durable feature here, more than the price cut. Being able to ask for less reasoning on a request you already know is simple means you can keep one model deployed and tune spend per call instead of maintaining two client configs. The catch is that effort and model choice solve different problems. Effort controls trade capability for tokens within a tier; they don't make Haiku 5.5 suddenly safe for high-stakes work at max effort, and they don't make Sonnet cheap at minimum effort. Treat effort as a dial you turn after you've already decided the tier, not a substitute for the decision.

Batch and committed capacity don't travel evenly#

Batch API savings are reportedly still around 50% off standard rates for Haiku 5.5, which stacks well with its already-low ceiling for anything that doesn't need a synchronous response: nightly summarization, backfill extraction, bulk classification. But Priority Tier, the committed-capacity option for guaranteed throughput at peak traffic, is reportedly not available for Haiku 5.5. If part of your routing logic exists because you need predictable latency under load, not just low cost, that's a reason to keep the request on a model tier that actually offers it rather than assuming Haiku covers every operational need Haiku 4.5 did.

The decision, concretely#

  • Is this a subagent call feeding a larger pipeline where a bad result gets caught downstream? Route to Haiku 5.5, and use effort controls to tune cost further.
  • Does the prompt regularly run near or past 100K tokens under the new tokenizer? Check the actual count before assuming the discount applies; past that line, the pricing gap shrinks a lot.
  • Is a wrong answer here one a user or a dollar amount will notice? Stay on Sonnet 5.5 or GPT-6.1 Sol, full stop.
  • Do you need guaranteed throughput at peak load, not just a low price? Priority Tier isn't on Haiku 5.5, so that requirement alone should keep the call on a frontier model.

The call we'd make#

Our default is to route high-volume, low-stakes, easily-retried work to Haiku 5.5 immediately, measure actual token counts on our own prompts rather than trusting the 75% headline, and leave everything else exactly where it was. If you haven't already picked a frontier default for the calls that stay there, our comparison of the current LLM API options is the place to settle that, separately from this decision.

Explore topics:AI
React

Get the DevOps Troubleshooting Cheat Sheet

Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.

Share this post
KU

Kiril Urbonas

AI Engineer

567 articles
View all articles by Kiril Urbonas

You might have missed

Evergreen posts worth revisiting.