Patman's Neural Network

Loading knowledge network

GPT-6.1 Sol: What matters is the completed task, not the token price

OpenAI promises performance close to Astra at significantly lower costs. Whether the switch is worth it depends on your own workflows. Published on 29 September 2026 • AI translated

A cheaper model saves little if the more expensive one has to finish the job afterward. That is precisely what GPT-6.1 Sol must be measured against: the cost per successfully completed task, not the price tag per token.

OpenAI positions GPT-6.1 Sol as a model for software engineering, computer operation, and professional applications. According to available information, it is available via the API and is being rolled out gradually in ChatGPT Work and Codex.

The reported test results are impressive. On DeepSWE v1.1, OpenAI states that Sol reaches the coding performance of GPT-6 Astra at roughly one-fifth of the task cost. In the offline test of OSWorld 2.0 at maximum reasoning effort, it trails Astra by 2.1 percentage points at about one-seventh of the cost per task.

These are results under specific test conditions, not a blanket guarantee of savings. Token consumption, reasoning settings, and the type of task affect the calculation. "Close to Astra" does not mean Sol replaces Astra everywhere.

Token type Reported price per million tokens
Input without cache 2 US dollars
Cached input 0.10 US dollars
Output 10 US dollars

The cache is particularly interesting. Coding agents and document workflows frequently resend the same instructions, code context, or documents. If these inputs are recognized as reusable, the lower rate applies.

At the stated prices, ten million cached input tokens cost 1 US dollar instead of 20 US dollars without caching. However, repeated content alone does not guarantee a cache hit. Uncached inputs and generated outputs remain at the regular rate.

OpenAI also reports progress regarding factual errors: at low reasoning effort, the rate of incorrect answers drops from 11.4 to 7.7 percent compared to GPT-6 Sol. The benchmark deliberately uses challenging conversations where users had previously flagged errors. This rate therefore does not describe normal operations.

For teams, this does not mean an automatic model switch, but rather a clean comparison. Sol and Astra must solve the same internal tasks using the same acceptance criteria. Four measurement points are enough to reveal the crucial differences.

1. Verify task success

Define in advance when a task counts as completed. Then measure how often each model actually meets these criteria.

2. Capture total costs

Add up inputs, outputs, and actual cache hits. Also factor in necessary fallbacks to Astra, rather than only counting the cheaper initial attempt.

3. Compare runtimes

Measure median and tail latencies across the required reasoning settings. Benchmark costs do not indicate how quickly your own workflow will complete.

4. Check boundaries

Check how Sol handles failing tools, explicit restrictions, and permissions. Unsubstantiated claims and unauthorized actions must also be factored into the evaluation.

Cheaper only matters when it works

The reported results make Sol a viable candidate for more cost-effective AI agents. However, the decision comes down to your own production environment: Does it complete the work reliably, and does it require Astra as a fallback rarely enough? Only then does a low token price turn into genuine savings.