GPT-6 Sol & Claude Opus 5.5 are liveGPT-6 Sol at half the price of 5.6 Sol
Blog
Guide

GPT-6.1 Sol review

Judge the upgrade by finished work and total cost.

10 min readOmniaKey
GPT-6.1 SolGPT-6 AstraCodingAPI pricingModel review

GPT-6.1 Sol is worth evaluating first if you already use GPT-6 Sol for complex coding and repeated revisions. Its standard input and output rates are unchanged, cache reads cost half as much, and OpenAI reports stronger results on several evaluations. Keep Astra in the comparison when mistakes are expensive or a task requires sustained investigation.

The important distinction is between a lower token rate and a cheaper completed task. OpenAI's own coding chart includes a case where new Sol costs slightly more per task than old Sol while scoring substantially higher. That is a more useful starting point than a blanket promise of “Astra for one-fifth the cost.”

Sources checked September 30, 2026 UTC. This is a review of official documentation and attributed benchmark results. We have not run an independent model comparison. Cost examples use assumed token quantities. OmniaKey publishes this article and provides an API gateway; its quotes are separate from OpenAI's direct API prices.

What changes with GPT-6.1 Sol?

OpenAI's API changelog records the release on September 29, 2026. The exact model ID is gpt-6.1-sol. Its stated focus is complex coding, computer use and professional work at a lower cost than Astra.

SpecificationGPT-6.1 Sol
Context window1,050,000 tokens
Maximum output128,000 tokens
Native input / outputText and images / text
API reasoning effortlow, medium (default), high, xhigh, max
Tool-calling APIResponses

The new Sol model page, old Sol page and Astra page list the same context and output capacities. This upgrade does not expand the window. Reasoning and output still occupy context; equal capacity does not establish equal retrieval accuracy. Native audio and video are not supported by this model.

Coding results: compare the same reasoning setting

The interactive DeepSWE v1.1 chart in OpenAI's announcement reports these results for complex software-engineering tasks. All three rows below use medium. The costs are the publisher's benchmark measurements, not our synthetic examples or OmniaKey bills.

Model at mediumDeepSWE v1.1 scoreAverage cost per task
GPT-6.1 Sol73.0%$0.42
GPT-6 Sol56.6%$0.38
GPT-6 Astra72.8%$3.08

New Sol approaches Astra's score at a much lower task cost in this evaluation. The 0.2-percentage-point gap does not prove that it is generally better than Astra, and no statistical significance is established here. Against old Sol, its task cost actually rises from $0.38 to $0.42 despite unchanged standard input/output rates. Token use matters.

More reasoning is not automatically better: new Sol reaches 75.2% at high, costing $0.65, but 71.9% at max, costing $1.57. The announcement's 6.4-point improvement compares new Sol at high with old Sol's best score of 68.8% at max; it is not the medium-to-medium difference.

Other reported improvements are relevant to agents: AutomationBench 1.0.6 gains 4.8 percentage points over old Sol at medium. On OSWorld 2.0's offline v2026.08.08 set, the max partial-reward score improves 7 points and trails Astra max by 2.1 points. Partial reward is not a universal desktop-task completion rate.

On deliberately difficult factuality prompts, answers containing at least one factual error fall from 11.4% to 7.7% at low. These selected conversations are not representative of everyday use. Astra also retains the highest score in the publisher's Terminal-Bench Science 0.1 comparison, 68.1% at max. The evidence supports evaluating new Sol, not retiring Astra from every workflow.

API pricing: what actually gets cheaper?

These are OpenAI direct Standard rates in USD per million tokens, for requests with at most 272K input tokens, checked against API pricing. Subscription allowances, gateway quotes, tools and taxes are separate.

Token categoryGPT-6.1 SolGPT-6 SolGPT-6 Astra
Uncached input$2.00$2.00$10.00
Cache read$0.10$0.20$1.00
Cache write$2.50$2.50$12.50
Output$10.00$10.00$50.00

Relative to old Sol, only cache reads fall among these four rates. Against Astra, standard input and output are one-fifth the price; reads are one-tenth. Neither ratio determines the cost of an entire task.

Four reproducible cost examples

K means 1,000 tokens. Output includes billed reasoning. These examples hold all models' token counts constant; actual runs can differ. A valid cache hit is assumed only in the read row, with its earlier write billed separately. Tools, infrastructure, taxes and gateway fees are excluded.

Assumed workloadGPT-6.1 SolGPT-6 SolGPT-6 Astra
100K uncached input + 10K output$0.30$0.30$1.50
200K initial cache write + 20K uncached input + 10K output$0.64$0.64$3.20
200K cache read + 20K uncached input + 10K output$0.16$0.18$0.90
300K uncached input + 10K output, long-context rates$1.35$1.35$6.75

For new Sol's read row, 0.20 × $0.10 + 0.02 × $2 + 0.01 × $10 = $0.16. Old Sol costs $0.18. A 50% cache-read discount saves only 11.1% on this request. Include one initial write and ten identical valid reads: $2.24 versus $2.44, a saving of approximately 8.2%. That assumes the prefix stays valid and every read hits.

The caching guide says cache-write pricing is not additive. Classify input as uncached, read or write; do not charge the same write tokens again as ordinary input.

The 272K threshold applies to the whole request

Above 272K input tokens, input and cache rates double and output rates rise by 1.5 times for the entire request. Cached input counts toward input length. The 300K example therefore costs 0.30 × $4 + 0.01 × $15 = $1.35 on new Sol.

Reasoning tokens are billed as output. Fast mode costs twice Standard; Batch and Flex cost half. Price multipliers do not establish end-to-end speed. Record retries and tool waits alongside token usage.

Three checks before upgrading from old Sol

The GPT-6 guide identifies request-level differences that can affect an existing agent.

  1. Reasoning: old Sol supports none; new Sol does not support none or minimal. Move to a supported setting and measure usage and latency again. The API default is medium; client defaults can differ.
  2. Tools: old Sol supports function calling through Chat Completions only with reasoning_effort: "none". New Sol requires Responses for tool calling. Chat Completions remains available for requests without tools.
  3. Sampling parameters: remove unsupported settings such as temperature and top_p when reasoning is enabled. Check what the SDK or gateway injects automatically.

The documented Codex CLI selection is:

bash
codex --model gpt-6.1-sol

Use /model inside an interactive session. These are documented instructions, not an account invocation tested for this review.

Where is it available?

The Codex model guide lists launch access for eligible Plus, Pro, Business, Enterprise and Edu surfaces in Work and Codex. Enterprise and Edu administrators must enable it; Free and Go are outside the launch rollout. In ChatGPT, it is available in Work and Codex rather than ordinary Chat at launch. Check your account, client and workspace for available speed modes.

Subscription and credit usage are separate from API prices. Codex credits have no separate cache-write charge. Do not convert the API examples above into a promised number of included subscription tasks.

Is it worth making GPT-6.1 Sol your default?

For old Sol users, the combination of unchanged standard rates and stronger published coding results makes a controlled upgrade trial worthwhile. For Astra users, start with frequent work that has clear acceptance checks; retain the stronger alternative when failures or human corrections dominate cost.

A practical pilot is six tasks: two reproduced bugs, two cross-file changes and two evidence-based document questions. Run each model three times per task with a fresh workspace, fixed inputs, tools, permissions and retry budget. Start all three at explicit medium, then report any tuned-effort experiment separately. This is a proposed test, not one we performed.

Record acceptance, omissions, total elapsed time, billed tokens, failed attempts and human corrections. Divide all attempt costs by accepted results; if none succeed, report no accepted result. This small pilot cannot establish a universal success rate.

For version-specific background, see the old GPT-6 Sol review and old Sol/Astra comparison. Check the current model catalog for OmniaKey's displayed routes and quotes. Official release status does not establish availability through a particular gateway. The OpenAI sources linked throughout are in English.