GPT-6.1 Sol's 80% Price Cut Rewrites Agentic Coding Math
OpenAI priced GPT-6.1 Sol at a fifth of GPT-6 Astra and called it nearly as good at agentic coding. That claim is now a routing decision, not a headline.
Key takeaways
- OpenAI priced GPT-6.1 Sol at a fifth of GPT-6 Astra and called it nearly as good at agentic coding.
- That claim is now a routing decision, not a headline.
On this page
OpenAI shipped GPT-6.1 Sol at DevDay on September 29 at $2 per million input tokens and $10 per million output tokens, a fifth of what GPT-6 Astra charges, and said it nearly matches Astra's intelligence on agentic coding, computer use, and professional work. If your coding agent, browser-automation pipeline, or internal copilot has been defaulting to Astra out of caution, that default just got a lot more expensive to justify. The right move is not to flip every workload to Sol overnight either. It is to treat this as a routing problem and let your own evals, not OpenAI's benchmark slide, decide where the line sits.
The cut, and why this one is real#
GPT-6 Astra launched on September 3 at $10 per million input tokens and $50 per million output tokens. Sol comes in at $2 and $10, with cached input at $0.10 per million tokens, which is 95% off the standard input rate. For anything over 272,000 input tokens, Sol moves to a long-context tier priced at $4 input and $15 output, with cached input at $0.20. That is still roughly a quarter of Astra's standard rate even in the expensive tier.
What makes this cut different from the usual "new model, slightly cheaper" pattern is the claim attached to it. OpenAI isn't positioning Sol as a budget option, but as near-Astra on the categories that drive spend: agentic coding, computer use, and professional work. On DeepSWE v1.1, a software-engineering benchmark run against real codebases, Sol reportedly matches Astra's score at around a fifth of the cost. If that holds on your codebase, every agent you run against Astra today pays a 5x markup for nothing.
Where Sol is good enough, and where Astra still wins#
"Nearly matches" is doing real work in that sentence, and it is worth taking literally rather than rounding up to "equivalent." Astra reportedly still leads on the hardest scientific and research-grade reasoning tasks, and OpenAI's own guidance points toward Astra for that narrow slice of work. The gap seems to concentrate at the top end of difficulty, not across the board.
Where Sol wins on cost without a quality trade you'll notice: day-to-day agentic coding tasks, routine computer-use flows (clicking through a web app, filling forms, running a known sequence of tool calls), and professional-work generation like drafting reports, specs, or structured summaries. These are high-volume, moderate-difficulty tasks where a fifth of the price is nearly a fifth of your bill with no measurable drop in output quality.
Where you probably still want Astra: multi-hop scientific reasoning, the hardest refactors across unfamiliar codebases, and anything where a wrong answer is expensive to catch downstream. If your task sits in that category, the 5x price difference is not the number that matters. The failure rate is.
The benchmark trap before your own evals confirm it#
Launch-day benchmarks are marketing with numbers attached. DeepSWE v1.1 and Terminal-Bench Science show how Sol performs on OpenAI's chosen tasks, not your test suite or tool-calling setup. A model that nearly matches Astra on a public benchmark can still regress on your agent loop if your tasks lean on capabilities the benchmark skips, like state tracking across fifteen tool calls or jargon in a legacy codebase.
Before moving a production workload from Astra to Sol, run your own eval set against both and compare pass rate, not vibes. No eval set yet? Build one from your last 50 real agent tasks first. A 20% miss rate on Sol that forces a retry on Astra erases the savings and adds latency on top.
Rewriting the routing layer#
Most teams already have a model-tier router from the Astra era that assumed "frontier" meant "Astra, always." That assumption is now wrong by default. A simple routing function that checks task class and context length, with an explicit escape hatch to the harder model, looks like this:
# route.py: pick a model tier post-Sol
def pick_model(task, input_tokens, eval_confidence):
# Escalate automatically past Sol's long-context tier threshold
long_context = input_tokens > 272_000
if task.category in ("agentic_coding", "computer_use", "professional_work"):
if eval_confidence.get(task.category, 0) >= 0.9:
return "gpt-6.1-sol-long" if long_context else "gpt-6.1-sol"
# Haven't validated this task class against your own evals yet
return "gpt-6-astra"
if task.category in ("science_reasoning", "hard_refactor", "high_stakes"):
return "gpt-6-astra"
return "gpt-6.1-sol-long" if long_context else "gpt-6.1-sol"
model = pick_model(task, input_tokens=18_000, eval_confidence=eval_scores)
The eval_confidence gate is the part teams skip under deadline pressure, and it is the part that keeps this safe. Route to Sol by default, but only after a task category has cleared your own bar, not OpenAI's.
The long-context tier matters more than it looks#
The 272,000-token threshold is not an edge case for agentic workloads. A coding agent loading a repo's file tree, several large source files, and a running conversation history crosses that line fast, past the 10th or 15th tool call in a session. At $4 input and $15 output, Sol's long-context tier is still cheap relative to Astra, but it is twice the price of Sol's standard tier, so sloppy context management now costs twice: once in token count, once in tier multiplier. Trim stale tool outputs and summarize old turns before a session crosses the threshold, the same discipline that matters for caching, covered in our best LLM APIs guide.
The decision, concretely#
- Is the task agentic coding, computer use, or professional-work generation, and has your eval suite cleared Sol on that category? Route to Sol and bank the 80% savings.
- Is the task multi-hop scientific reasoning or a high-stakes refactor where a wrong answer is costly? Keep it on Astra until you have hard evidence Sol closes that specific gap.
- Is a session regularly crossing 272,000 input tokens? Fix context bloat first, since the long-context tier doubles Sol's price before you even compare it to Astra.
- Have you actually run your own benchmark, or are you trusting OpenAI's DeepSWE v1.1 number? If it's the latter, you're routing on marketing, not data.
The call we'd make#
Default new agentic coding, computer-use, and professional-work traffic to GPT-6.1 Sol, gated behind a same-day eval run against your last month of real tasks, not the DevDay slides. Keep Astra as the fallback for anything scientific or high-stakes until your numbers say otherwise, and watch the 272,000-token boundary as closely as the per-token price, since a bloated context window is where this price cut quietly disappears. If your billing already tracks model-tier spend the way we described for Copilot's credit system, this is the same lever: the model picker, not the vendor, moves the invoice.
Get the DevOps Troubleshooting Cheat Sheet
Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.
Azure OpenAI's Sweden Central Outage: A Health Check Postmortem
A slow database made a backend service look dead, and the health check that was supposed to protect it killed it faster. The lesson has nothing to do with AI.
Gemini 4 Argon: A Frontier Model You Mostly Can't Use Yet
Google shipped a 1M-token output ceiling and strong cyber-defense benchmarks, then handed the keys to almost nobody. Here is what that gating pattern tells you.
More from AI
Explore more articles in this category
Claude Code Mods: Programmable, Unsandboxed, and Your Problem in CI
Mods make Claude Code a real runtime you can program. They also run with your full native permissions, no sandbox, so a bad one is a bad CI step.
Gemini 4 Argon: A Frontier Model You Mostly Can't Use Yet
Google shipped a 1M-token output ceiling and strong cyber-defense benchmarks, then handed the keys to almost nobody. Here is what that gating pattern tells you.
Claude Sonnet 5.5: Should You Switch From Sonnet 5
Same sticker price, a real-world cost drop from speed and token efficiency. Here is who should move today and who can wait a sprint.
You might have missed
Evergreen posts worth revisiting.