GPT-6.1 Sol API Guide (2026): Pricing, Migration From GPT-6 Sol, and Astra Comparison

OpenAI released GPT-6.1 Sol at DevDay on September 29, 2026, one week after GPT-6 Sol went out. The pitch is simple: near-Astra quality for agentic coding, computer use, and professional work at one-fifth of GPT-6 Astra's per-token price. TechCrunch also reported that OpenAI isn't shipping a GPT-6.1 Astra, so the new Sol is the OpenAI model most teams will build against this quarter.

The list price matches GPT-6 Sol, so you'd expect a one-line model swap. Mostly, it is. But cache pricing changed, two request patterns that work on GPT-6 Sol stop working, and the 272K long-context surcharge still catches people off guard. This guide covers all of it, with code.

Key takeaways

  • OpenAI released GPT-6.1 Sol (API model ID gpt-6.1-sol) at DevDay on September 29, 2026, one week after GPT-6 Sol.
  • GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens, the same as GPT-6 Sol, while GPT-6 Astra costs $10 input and $50 output.
  • GPT-6.1 Sol bills cached input at $0.10 per million tokens (5% of the input price), half of GPT-6 Sol's $0.20 cached rate.
  • GPT-6.1 Sol has a 1,050,000-token context window, 922,000 max input tokens, and 128,000 max output tokens; prompts above 272,000 input tokens are billed at 2x input and 1.5x output for the whole request.
  • GPT-6.1 Sol does not support the none or minimal reasoning efforts, and its tool calling requires the Responses API.

What OpenAI Shipped on September 29

GPT-6.1 Sol takes over the middle of OpenAI's lineup. The GPT-6 Sol model page already points readers to GPT-6.1 Sol as "the newer Sol model." OpenAI says the update is better at programming and debugging, document understanding, and multi-step workflows, and that on several of those it gets close to GPT-6 Astra.

The one hard quality number is factuality. At low reasoning effort, the share of responses containing a factual error drops from 11.4% on GPT-6 Sol to 7.7% on GPT-6.1 Sol, and OpenAI says the error rate stays within 1.9 points of GPT-6 Astra across reasoning settings. Those are OpenAI's own evals. Treat them as a hypothesis for your workload, not a result.

GPT-6.1 Sol API Pricing

Prices below are OpenAI's and Anthropic's published Standard-tier list prices as of October 1, 2026, in USD per 1M tokens.

ModelInputCached inputOutputContext windowAbove 272K input
GPT-6.1 Sol$2.00$0.10$10.001,050,000$4.00 in / $15.00 out
GPT-6 Sol$2.00$0.20$10.001,050,000$4.00 in / $15.00 out
GPT-6 Astra$10.00$1.00$50.001,050,000$20.00 in / $75.00 out
Claude Sonnet 5.5$2.00$0.20$10.001,000,000No surcharge ($2.00 / $10.00)

Batch and Flex are 50% off, which puts GPT-6.1 Sol at $1 input and $5 output. Fast mode is 2x Standard ($4 / $20). Cache writes cost $2.50 per 1M tokens (1.25x input), and regional processing adds 10% where it's offered. Ultrafast mode is listed only for GPT-6 Astra, at $60 input and $300 output.

Where the bill actually changes

Same sticker price, so where's the saving? Cached input. GPT-6.1 Sol bills cache hits at 5% of input instead of 10%. Agents live on cache hits, because every turn re-sends the repo map, tool schemas, and the conversation so far.

Take a 30-turn coding session where each turn sends 120K cached tokens and 8K fresh input tokens, and gets 2K output tokens back (ignoring the one-time cache write):

ModelCost per turn30-turn session
GPT-6.1 Sol$0.048$1.44
GPT-6 Sol$0.060$1.80
GPT-6 Astra$0.300$9.00

That's 20% cheaper than GPT-6 Sol and 84% cheaper than Astra on identical traffic. Plug in your own numbers with the API cost calculator.

Watch the 272K line

Once a prompt goes over 272,000 input tokens, the whole request is billed at 2x input and cache rates and 1.5x output. Not just the overflow. A 270K-token prompt costs $0.54 in uncached input on GPT-6.1 Sol. A 280K-token prompt costs $1.12. Agent transcripts drift past that line quietly, so compact before you get there (code below).

GPT-6.1 Sol vs GPT-6 Sol vs GPT-6 Astra

AttributeGPT-6.1 SolGPT-6 SolGPT-6 Astra
Model IDgpt-6.1-solgpt-6-solgpt-6-astra
Input / output per 1M$2 / $10$2 / $10$10 / $50
Cached input per 1M$0.10 (5% of input)$0.20 (10% of input)$1.00 (10% of input)
Context / max output1,050,000 / 128,0001,050,000 / 128,0001,050,000 / 128,000
Reasoning effortlow, medium (default), high, xhigh, maxnone, low, medium (default), high, xhigh, maxlow, medium, high, xhigh, max
Tool calls in Chat CompletionsNo, Responses API onlyOnly with reasoning_effort: "none"No, Responses API only
Knowledge cutoffApril 30, 2026April 20, 2026April 30, 2026
Service tiersStandard, Batch, Flex, FastStandard, Batch, Flex, FastStandard, Batch, Flex, Fast, Ultrafast
Best forDefault agentic coding, computer use, long cached agent loopsExisting apps that rely on effort none for fast tool callsThe hardest tasks, where quality matters more than cost
Key limitationNo none/minimal effort; no audio, video, or fine-tuningSuperseded by GPT-6.1 Sol; cache hits cost 2x more5x the per-token price of either Sol model

Migrating From GPT-6 Sol: What Breaks

Changing the model string is the easy part. Check these three things first.

  1. No none or minimal effort. GPT-6.1 Sol accepts low, medium (default), high, xhigh, and max. OpenAI's advice is to use low where you had none, and to start at low if you used minimal.
  2. Tool calls need the Responses API. On GPT-6 Sol you could call functions through Chat Completions with reasoning_effort: "none". GPT-6.1 Sol has no none, and it supports Chat Completions only for requests without tools. Any Chat Completions plus tools path has to move to /v1/responses.
  3. Drop sampling parameters. When effort isn't none, OpenAI says to remove temperature, top_p, and top_logprobs (plus logprobs in Chat Completions). On GPT-6.1 Sol, that means every request.

Two smaller notes. If you change effort mid-conversation, use configuration_update items and keep the request-level reasoning.effort fixed so the cached prefix survives. And if you're jumping straight from GPT-5.5 or earlier, replace prompt_cache_retention with prompt_cache_options.ttl set to "30m".

curl: first request

curl https://api.openai.com/v1/responses \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6.1-sol",
    "reasoning": {"effort": "medium"},
    "input": "Find the race condition in this Go worker pool and propose a minimal fix: ..."
  }'

Python: a migration shim

This wraps old GPT-6 Sol request params so they're safe for GPT-6.1 Sol, then logs cache hits so you can watch the 5% rate do its job.

from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from the environment

# GPT-6.1 Sol doesn't support these efforts; map them before switching.
EFFORT_MAP = {"none": "low", "minimal": "low"}
SAMPLING_KEYS = ("temperature", "top_p", "top_logprobs")


def to_gpt61_sol(params: dict) -> dict:
    req = dict(params)
    req["model"] = "gpt-6.1-sol"
    effort = req.get("reasoning", {}).get("effort", "medium")
    req["reasoning"] = {"effort": EFFORT_MAP.get(effort, effort)}
    for key in SAMPLING_KEYS:
        req.pop(key, None)  # not allowed when effort is not "none"
    return req


resp = client.responses.create(**to_gpt61_sol({
    "input": "Summarize the failing tests and suggest the smallest fix.",
    "reasoning": {"effort": "none"},  # legacy GPT-6 Sol setting
    "temperature": 0.2,
}))

usage = resp.usage
cached = usage.input_tokens_details.cached_tokens
print(resp.output_text)
print(f"input={usage.input_tokens} cached={cached} output={usage.output_tokens}")

Node.js: compact before the 272K line

A cheap guard: estimate tokens and summarize the transcript before it crosses the surcharge threshold. The same code works against any OpenAI-compatible endpoint by changing baseURL. If you route through KissAPI or another gateway, confirm it exposes gpt-6.1-sol and the Responses endpoint before you move production traffic.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.API_KEY,
  baseURL: process.env.API_BASE_URL, // e.g. https://api.openai.com/v1
});

// The surcharge starts above 272K input tokens; compact with headroom.
const COMPACT_AT = 250_000;

// Rough estimate: about 4 characters per token for English and code.
const estimateTokens = (text) => Math.ceil(text.length / 4);

export async function runTurn(history, userMsg) {
  let context = history.join("\n\n");

  if (estimateTokens(context) > COMPACT_AT) {
    const summary = await client.responses.create({
      model: "gpt-6.1-sol",
      reasoning: { effort: "low" },
      input: `Compress this agent transcript into the facts, decisions, and open TODOs needed to continue:\n\n${context}`,
    });
    context = summary.output_text;
  }

  return client.responses.create({
    model: "gpt-6.1-sol",
    reasoning: { effort: "medium" },
    input: `${context}\n\n${userMsg}`,
  });
}

Compaction rewrites your prefix, so the next turn pays one cache write. That's still far cheaper than paying 2x on every turn above 272K. For better estimates than four characters per token, paste real prompts into the token counter.

Which Model Should Get Your Traffic?

My take: make GPT-6.1 Sol the default this week, but don't delete the Astra route. "Near-Astra" is an average across benchmarks, and the hardest 5% of your tasks is exactly where an average hides the gap. Pull 50 to 100 real tasks from your logs, run them on both models, and compare cost per passed task rather than cost per token. A model that needs a second attempt costs twice as much, whatever its rate card says.

FAQ

How much does the GPT-6.1 Sol API cost?

GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens. Cached input costs $0.10 per million tokens and cache writes cost $2.50. Batch and Flex are 50% off at $1 input and $5 output. Prompts above 272,000 input tokens are billed at 2x input and 1.5x output for the full request ($4 / $15).

What is the difference between GPT-6.1 Sol and GPT-6 Sol?

Both cost $2 input and $10 output per million tokens with a 1,050,000-token context window. GPT-6.1 Sol halves cached input to $0.10, has an April 30, 2026 knowledge cutoff, and scores closer to GPT-6 Astra on OpenAI's evals. It drops the none and minimal efforts, and its tool calling requires the Responses API.

Is GPT-6.1 Sol as good as GPT-6 Astra?

OpenAI says it's near-Astra for agentic coding, computer use, and professional work at one-fifth of Astra's per-token price ($2 / $10 vs $10 / $50). Those are OpenAI's own evals, and GPT-6 Astra is still OpenAI's most capable model. Test your hardest tasks on both before moving traffic.

Compare OpenAI and Claude Models With One Key

Create a free KissAPI account and A/B test models through one OpenAI-compatible endpoint before you commit production traffic.

Start Free