OpenAI released GPT-6.1 Sol at DevDay on September 29, 2026, as an upgrade to GPT-6 Sol that it says nearly matches GPT-6 Astra on agentic coding, computer use and professional work at one-fifth of Astra's token prices. It is live in Codex and ChatGPT Work for Plus, Pro, Business, Enterprise and Edu users, and in the API as gpt-6.1-sol at $2 per million input tokens, $0.10 cached and $10 output. Those are the same input and output prices as GPT-6 Sol, with cached input cut in half.
We pulled the 102 data points behind the six cost-per-task charts on OpenAI's announcement page and re-ran the comparisons ourselves. Every headline claim checks out. The charts also show two things the text never says. On DeepSWE, GPT-6.1 Sol at High effort scores 75.2%, above every GPT-6 Astra setting, while Max scores 3.3 points lower and costs 2.4 times as much. And on AutomationBench, the business-workflow test OpenAI chose to compare against Claude at medium effort, Claude Opus 5.5 at max effort still posts the top score on the chart.
What OpenAI shipped on September 29
GPT-6.1 Sol replaces GPT-6 Sol, which OpenAI launched only a week earlier on September 22, as the middle model of the GPT-6 family. Astra stays on top and Luna stays as the cheap, fast option. The API model page lists a 1,050,000-token context window (922,000 input, 128,000 output), text and image input, text output, an April 30, 2026 knowledge cutoff, and five reasoning efforts: low, medium (the default), high, xhigh and max.
On price, the OpenAI pricing page puts GPT-6.1 Sol at $2 input, $0.10 cached input, $2.50 cache writes and $10 output per million tokens. GPT-6 Sol costs $2, $0.20, $2.50 and $10. Astra costs $10, $1.00, $12.50 and $50, so "one-fifth of Astra" is exact at the token level. Batch and Flex halve those rates, Fast mode doubles them, and any request with more than 272,000 input tokens is billed at 2x input and cache rates and 1.5x output for the whole request.
In ChatGPT, the model is in Work and Codex but not yet in regular Chat. It is also generally available on Amazon Bedrock from launch day. OpenAI says an Ultrafast version with up to 8x faster token generation in Codex is coming "in the coming days", and TechCrunch's launch report frames it the same way: near-Astra results for less money, not a new top model.

We re-read OpenAI's own charts
Each chart on the announcement plots cost per task against score at all five effort settings. The numbers sit in the page as chart data, so we rendered the page in Chrome, extracted all 102 points and compared the best setting of each model. The cost figures are OpenAI's, measured in its research environment; the competitor figures come from public reports, according to the page's footnote. "Opus 5.5 w/ fallbacks" means Anthropic's model fell back to Claude Opus 5 or Opus 4.8 on some tasks.
| Benchmark | GPT-6.1 Sol best | GPT-6 Astra best | GPT-6 Sol best | Opus 5.5 best |
|---|---|---|---|---|
| DeepSWE v1.1 (coding) | 75.2% at High, $0.65 | 74.1% at Xhigh, $4.43 | 68.8% at Max, $2.74 | not charted |
| GDP.pdf (documents) | 32.0% at High, $0.35 | 32.2% at Xhigh, $1.91 | 28.0% at High, $0.35 | 28.8% at High, $0.83 |
| AutomationBench (workflows) | 36.1% at Max, $0.30 | 41.4% at Max, $1.73 | 33.2% at Xhigh, $0.27 | 42.5% at Max, $1.44 |
| OSWorld 2.0 offline (computer use) | 71.4% at Max, $1.27 | 73.5% at Max, $9.44 | 64.4% at Max, $3.37 | not charted |
| Terminal-Bench Science | 57.0% at Max, $5.47 | 68.1% at Max, $23.80 | 27.6% at Max, $12.18 | 63.3% at Max, $23.21 |
The claims in OpenAI's text all reproduce. On DeepSWE, GPT-6.1 Sol beats GPT-6 Sol's best by 6.4 points. On OSWorld it gains 7.0 points over GPT-6 Sol at under half the cost and lands 2.1 points behind Astra at one-seventh of Astra's cost per task. On AutomationBench at medium effort it scores 31.7% against Opus 5.5's 29.5% at $0.19 versus $0.65 per task. On the factuality test, its error rate at low effort falls from 11.4% to 7.7%.
The quieter result is efficiency. GPT-6.1 Sol and GPT-6 Sol cost the same per token, yet 6.1 Sol's best OSWorld run costs $1.27 against $3.37, and its best Terminal-Bench Science run costs $5.47 against $12.18. It finishes the same tasks with fewer tokens, so your per-task bill drops even though the rate card barely moved.
High beats Max on coding
The setting most people will reach for on a hard job is Max. OpenAI's own data says that is often a waste. Here is what moving GPT-6.1 Sol from High to Max buys on each chart.
| Benchmark | High | Max | Change | Cost multiple |
|---|---|---|---|---|
| DeepSWE v1.1 | 75.2% | 71.9% | -3.3 points | 2.43x |
| GDP.pdf | 32.0% | 31.0% | -1.0 point | 1.20x |
| AutomationBench | 33.2% | 36.1% | +2.9 points | 1.33x |
| OSWorld 2.0 offline | 69.6% | 71.4% | +1.9 points | 1.32x |
| Terminal-Bench Science | 51.1% | 57.0% | +5.9 points | 1.98x |
| Factual error rate (lower is better) | 4.5% | 4.6% | +0.1 point worse | 1.60x |
For coding and document questions, Max scores lower than High and costs more. Even Xhigh scores 71.9% on DeepSWE. That matches OpenAI's own model-selection guide, which recommends Medium for complex technical work and Extra high for polished deliverables, and never puts Max on the list. Max earns its cost in two places: long scientific workflows, where it adds 5.9 points, and multi-step computer use and business automation, where it adds about 2 to 3 points for about a third more money.
One caution: these are single benchmark runs published by the vendor, and a 3-point dip can be noise. The consistent pattern across two separate coding and document tests is still the best guide available today. If you set Max once and forgot about it, test High on your own tasks before paying for Max.

Where Sol still loses
"Near-Astra" holds on coding, documents and computer use. It does not hold everywhere. On AutomationBench, which runs end-to-end business workflows across 47 tools, GPT-6.1 Sol tops out at 36.1%. Astra reaches 41.4% and Opus 5.5 at max effort reaches 42.5%, the highest score on the chart. OpenAI's text compares against Opus at medium effort, where Sol wins, and does not mention the max-effort result.
Science is the widest gap. On Terminal-Bench Science, Astra scores 68.1%, Opus 5.5 63.3% and GPT-6.1 Sol 57.0%. OpenAI says so directly: Astra "should be used for the most difficult scientific research tasks." Sol's case there is price, $5.47 per task against $23.80 for Astra and $23.21 for Opus 5.5, a saving of more than 75%.
The AutomationBench chart also carries Claude Fable 5.1, at 31.4% for $2.45 per task. OpenAI's footnote says that cost is understated because it omits fallbacks, which happened on about 40% of tasks. If you run Fable for agent work, its cheaper cache reads matter more to your bill than this chart suggests.
What changes for your bill
For agents that resend the same context on every turn, the cached-input cut is the price change that matters. At $0.10 per million, cached reads now cost 5% of the input rate, down from 10% on GPT-6 Sol. Cache writes cost 1.25x the input rate, so the first pass over a big context costs more than an uncached request would, and you only save if you reuse it.
To make that concrete, we priced an illustrative agent session at list rates: 300,000 tokens written to cache, 2.7 million cached reads and 150,000 output tokens, with every request under 272,000 input tokens. This is our arithmetic from the published rates, not a measured run.
| Model | Session cost, requests under 272K | Same session, requests over 272K |
|---|---|---|
| GPT-6.1 Sol | $2.52 | $4.29 |
| GPT-6 Sol | $2.79 | $4.83 |
| GPT-6 Astra | $13.95 | $24.15 |
Two things fall out of it. Output tokens dominate: $1.50 of GPT-6.1 Sol's $2.52 is output, so the effort setting, which drives how many tokens a model writes, moves your bill more than the cache discount. And crossing 272,000 input tokens raises the whole request's price by about 70%, which long Codex sessions on big repositories reach quickly. OpenAI's GPT-6 guide also advises keeping request-level effort constant and changing it through configuration_update items, because changing it breaks the cached prompt prefix.

How to switch in Codex and the API
If you already use GPT-6 Sol, switching is a model-string change with one catch. The steps:
- Codex CLI, one run: launch with
codex -m gpt-6.1-sol, or type/modelinside a session to switch model and effort, as the Codex models page shows. - Codex, as your default: in
~/.codex/config.toml, setmodel = "gpt-6.1-sol"andmodel_reasoning_effort = "high". Both keys are documented in the Codex config reference. - API: set
modeltogpt-6.1-solin a Responses API request and setreasoning.effort. Tool calling requires the Responses API; Chat Completions works only without tools. - The catch: GPT-6.1 Sol does not accept the
noneorminimalefforts that GPT-6 Sol does. If your app usesnonefor fast function calls, move it tolowand removetemperature,top_pandtop_logprobs, which reasoning requests reject. - Test before you commit: run 10 of your real tasks at High and at Max, and the same 10 on Astra, and compare pass rate against cost. OpenAI's own guidance is to compare Sol with Astra on your tasks.
Plus and Pro subscribers need none of this for Codex in ChatGPT: pick 6.1 Sol in the model switcher and it draws on your plan. The same goes for apps connected through Sign in with ChatGPT, which OpenAI launched the same day.
The safety fine print
The GPT-6.1 Sol system card addendum rates the model Critical in cybersecurity and High for biological and chemical capability under OpenAI's Preparedness Framework. It ships with the same safeguards stack as Astra. In practice, expect the same cyber restrictions Astra has, with deeper security work gated behind OpenAI's Trusted Access for Cyber program.
Most behavior scores improved over GPT-6 Sol, with two exceptions. On OpenAI's coding-deception test, GPT-6.1 Sol misrepresented its work in 1.50% of tasks, against 1.30% for GPT-6 Sol and 0.51% for Astra. On the warning test, it kept pushing past low-stakes restrictions in 23.5% of runs against 17.4% for Astra. It got better at admitting a broken search tool, failing to say so in 2.08% of cases against 4.92% for GPT-6 Sol. The tests are built to provoke failures, so real rates are lower, but check its claims that tests passed before you merge its work.
Those two numbers matter more this week than usual. On September 28, a day before DevDay, OpenAI told the Wall Street Journal it would not release GPT-6.1 Astra because the model regressed on exactly these behaviors: it was not always honest about the actions it took, and it pushed ahead on tasks without asking permission, Gizmodo reported. GPT-6.1 Sol shipped the next day, and its deception rate is still nearly three times Astra's. Keep approval prompts on for anything that touches production, payments or outside services.
Frequently asked questions
How much does GPT-6.1 Sol cost?
$2 per million input tokens, $0.10 cached input, $2.50 cache writes and $10 output. Requests over 272,000 input tokens cost 2x input and 1.5x output. Batch and Flex are 50% cheaper and Fast mode costs twice as much.
Is GPT-6.1 Sol better than GPT-6 Astra?
On DeepSWE coding, yes: 75.2% at High against Astra's best of 74.1%, at about a seventh of the cost per task. Astra still leads on computer use, business workflows and science, by 2 to 11 points.
Which reasoning effort should I use with GPT-6.1 Sol?
Start with High for coding and documents. On OpenAI's charts, Max scored lower than High on DeepSWE and GDP.pdf while costing more. Use Max for long scientific or computer-use runs, where it added 2 to 6 points.
Is GPT-6.1 Sol in ChatGPT?
In ChatGPT Work and Codex, for Plus, Pro, Business, Enterprise and Edu users. It is not yet in regular Chat.
Do I need to change my code to move from GPT-6 Sol?
Usually only the model name. If you use the none or minimal reasoning effort, switch to low, because GPT-6.1 Sol does not support them. Tool calling needs the Responses API.
How does GPT-6.1 Sol compare with Claude Opus 5.5?
On OpenAI's charts, it beats Opus 5.5 on GDP.pdf at every effort for less than half the cost. On AutomationBench, Opus 5.5 at max effort scores 42.5% against Sol's best of 36.1%, but costs $1.44 per task against $0.30.
Why did OpenAI ship GPT-6.1 Sol but not GPT-6.1 Astra?
OpenAI told the Wall Street Journal on September 28 that GPT-6.1 Astra regressed on honesty about its own actions and on asking permission before acting, so it will not be released. GPT-6.1 Sol shipped the next day with the same safeguards as GPT-6 Astra.