Skip to main content

GPT-6.1 Sol: API Pricing, Long-Context Costs and a Migration Test

4 min read

GPT-6.1 Sol is in the API, ChatGPT Work and Codex, but not Chat. Here is what its token rates mean for cached and long-context jobs, plus a migration test.

GPT-6.1 Sol: API Pricing, Long-Context Costs and a Migration Test

A cheaper input price is easy to spot. The harder question is what your agent pays to finish one reliable task, especially when it reuses a large prompt or crosses the long-context threshold.

OpenAI introduced GPT-6.1 Sol on September 29, 2026. It is available through the API as gpt-6.1-sol and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. OpenAI says it is not yet available in Chat. That distinction matters if you are planning a migration for a chat workflow rather than an API or coding-agent workflow.

The price card has four relevant numbers.

The API model page lists standard rates per million tokens: $2 for uncached input, $0.10 for cached input, $2.50 for cache writes, and $10 for output. The model has a 1,050,000-token context window and a 128,000-token maximum output. A large window is capacity, not a reason to send every document on every call.

For a request with 40,000 uncached input tokens, 100,000 cached input tokens, and 10,000 output tokens, the token bill at standard rates is $0.19: $0.08 for new input, $0.01 for cached input, and $0.10 for output. That example assumes the cached prefix already exists. Writing 100,000 tokens to cache would add $0.25 at the listed cache-write rate. Tool calls and other charges are not included in this illustration.

Watch the 272,000-token breakpoint.

OpenAI says prompts with more than 272,000 input tokens are billed at twice the input and cache rates and 1.5 times the output rate for the full request. A 300,000-token uncached prompt with 20,000 output tokens therefore costs $1.50 before tool charges: $1.20 in input and $0.30 in output. It would be wrong to apply the higher rate only to the tokens above 272,000.

This changes how I would build retrieval. If a search system can give the model the few passages it needs instead of a 300,000-token document dump, the saving may come from crossing back under the threshold, not merely from shaving a few tokens. Measure any quality loss on real queries before shortening context by rule.

Don’t treat a benchmark result as a production promise.

OpenAI reports gains on coding, document analysis, and computer-use evaluations, including results near GPT-6 Astra on some tests. Those are company-reported measurements under stated settings, not a guarantee that Sol will finish your jobs at one-fifth the cost. The launch post also notes that research-environment evaluations can differ from production ChatGPT because prompts, tools, and effort settings differ.

The system-card addendum is worth reading for agent deployments. Its safety findings describe OpenAI’s own evaluation conditions; they do not replace your authorization checks, tool logs, or review of high-impact actions. Our guide to proof of presence for agent actions explains why a capable model and a permission to act are different decisions.

A migration test I would actually run.

  1. Choose 30 to 50 recent tasks, including easy cases, long-context cases,s and failures from the current model.
  2. Run the same tools, prompts, and success checks with the current model and GPT-6.1. Record reasoning effort and any fallback model.
  3. Log uncached input, cache writes, cache reads, output, tool charges, elapsed time, and whether a person had to repair the result.
  4. Compare cost per accepted task, not price per million tokens or number of generated lines.
  5. Roll out to a small share of traffic with a pinned model ID and a fast rollback path.

OpenAI lists low, medium, high, xhigh and max reasoning effort for this model; none and minimal are not supported. It also says tool calling belongs in the Responses API, while Chat Completions supports the model without tool calling. Check those differences before switching a production endpoint.

The decision

GPT-6.1 Sol is worth a controlled test if your current agent spends heavily on repeated context or on a more expensive model for ordinary coding and document work. I would hold off on a blind replacement where a missed action is costly. Our Sonnet 5.5 migration guide uses the same practical lens: actual task outcomes and full workflow cost should decide.

Read the source records.

Leave a comment

Your email address will not be published. Required fields are marked *