GPT-6.1 Sol API Guide (2026): Pricing, Migration From GPT-6 Sol, and Astra Comparison
OpenAI released GPT-6.1 Sol at DevDay on September 29, 2026, one week after GPT-6 Sol went out. The pitch is simple: near-Astra quality for agentic coding, computer use, and professional work at one-fifth of GPT-6 Astra's per-token price. TechCrunch also reported that OpenAI isn't shipping a GPT-6.1 Astra, so the new Sol is the OpenAI model most teams will build against this quarter.
The list price matches GPT-6 Sol, so you'd expect a one-line model swap. Mostly, it is. But cache pricing changed, two request patterns that work on GPT-6 Sol stop working, and the 272K long-context surcharge still catches people off guard. This guide covers all of it, with code.
Key takeaways
- OpenAI released GPT-6.1 Sol (API model ID
gpt-6.1-sol) at DevDay on September 29, 2026, one week after GPT-6 Sol. - GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens, the same as GPT-6 Sol, while GPT-6 Astra costs $10 input and $50 output.
- GPT-6.1 Sol bills cached input at $0.10 per million tokens (5% of the input price), half of GPT-6 Sol's $0.20 cached rate.
- GPT-6.1 Sol has a 1,050,000-token context window, 922,000 max input tokens, and 128,000 max output tokens; prompts above 272,000 input tokens are billed at 2x input and 1.5x output for the whole request.
- GPT-6.1 Sol does not support the
noneorminimalreasoning efforts, and its tool calling requires the Responses API.
What OpenAI Shipped on September 29
GPT-6.1 Sol takes over the middle of OpenAI's lineup. The GPT-6 Sol model page already points readers to GPT-6.1 Sol as "the newer Sol model." OpenAI says the update is better at programming and debugging, document understanding, and multi-step workflows, and that on several of those it gets close to GPT-6 Astra.
The one hard quality number is factuality. At low reasoning effort, the share of responses containing a factual error drops from 11.4% on GPT-6 Sol to 7.7% on GPT-6.1 Sol, and OpenAI says the error rate stays within 1.9 points of GPT-6 Astra across reasoning settings. Those are OpenAI's own evals. Treat them as a hypothesis for your workload, not a result.
- API model ID
gpt-6.1-sol, served on the Responses API, Chat Completions, and Batch. - In ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. Not yet in regular Chat.
- Text and image input, text output. No audio, video, or fine-tuning.
- Knowledge cutoff April 30, 2026 (GPT-6 Sol: April 20, 2026).
- US and EU data residency. Fast mode isn't available with EU residency.
GPT-6.1 Sol API Pricing
Prices below are OpenAI's and Anthropic's published Standard-tier list prices as of October 1, 2026, in USD per 1M tokens.
| Model | Input | Cached input | Output | Context window | Above 272K input |
|---|---|---|---|---|---|
| GPT-6.1 Sol | $2.00 | $0.10 | $10.00 | 1,050,000 | $4.00 in / $15.00 out |
| GPT-6 Sol | $2.00 | $0.20 | $10.00 | 1,050,000 | $4.00 in / $15.00 out |
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | 1,050,000 | $20.00 in / $75.00 out |
| Claude Sonnet 5.5 | $2.00 | $0.20 | $10.00 | 1,000,000 | No surcharge ($2.00 / $10.00) |
Batch and Flex are 50% off, which puts GPT-6.1 Sol at $1 input and $5 output. Fast mode is 2x Standard ($4 / $20). Cache writes cost $2.50 per 1M tokens (1.25x input), and regional processing adds 10% where it's offered. Ultrafast mode is listed only for GPT-6 Astra, at $60 input and $300 output.
Where the bill actually changes
Same sticker price, so where's the saving? Cached input. GPT-6.1 Sol bills cache hits at 5% of input instead of 10%. Agents live on cache hits, because every turn re-sends the repo map, tool schemas, and the conversation so far.
Take a 30-turn coding session where each turn sends 120K cached tokens and 8K fresh input tokens, and gets 2K output tokens back (ignoring the one-time cache write):
| Model | Cost per turn | 30-turn session |
|---|---|---|
| GPT-6.1 Sol | $0.048 | $1.44 |
| GPT-6 Sol | $0.060 | $1.80 |
| GPT-6 Astra | $0.300 | $9.00 |
That's 20% cheaper than GPT-6 Sol and 84% cheaper than Astra on identical traffic. Plug in your own numbers with the API cost calculator.
Watch the 272K line
Once a prompt goes over 272,000 input tokens, the whole request is billed at 2x input and cache rates and 1.5x output. Not just the overflow. A 270K-token prompt costs $0.54 in uncached input on GPT-6.1 Sol. A 280K-token prompt costs $1.12. Agent transcripts drift past that line quietly, so compact before you get there (code below).
GPT-6.1 Sol vs GPT-6 Sol vs GPT-6 Astra
| Attribute | GPT-6.1 Sol | GPT-6 Sol | GPT-6 Astra |
|---|---|---|---|
| Model ID | gpt-6.1-sol | gpt-6-sol | gpt-6-astra |
| Input / output per 1M | $2 / $10 | $2 / $10 | $10 / $50 |
| Cached input per 1M | $0.10 (5% of input) | $0.20 (10% of input) | $1.00 (10% of input) |
| Context / max output | 1,050,000 / 128,000 | 1,050,000 / 128,000 | 1,050,000 / 128,000 |
| Reasoning effort | low, medium (default), high, xhigh, max | none, low, medium (default), high, xhigh, max | low, medium, high, xhigh, max |
| Tool calls in Chat Completions | No, Responses API only | Only with reasoning_effort: "none" | No, Responses API only |
| Knowledge cutoff | April 30, 2026 | April 20, 2026 | April 30, 2026 |
| Service tiers | Standard, Batch, Flex, Fast | Standard, Batch, Flex, Fast | Standard, Batch, Flex, Fast, Ultrafast |
| Best for | Default agentic coding, computer use, long cached agent loops | Existing apps that rely on effort none for fast tool calls | The hardest tasks, where quality matters more than cost |
| Key limitation | No none/minimal effort; no audio, video, or fine-tuning | Superseded by GPT-6.1 Sol; cache hits cost 2x more | 5x the per-token price of either Sol model |
Migrating From GPT-6 Sol: What Breaks
Changing the model string is the easy part. Check these three things first.
- No
noneorminimaleffort. GPT-6.1 Sol acceptslow,medium(default),high,xhigh, andmax. OpenAI's advice is to uselowwhere you hadnone, and to start atlowif you usedminimal. - Tool calls need the Responses API. On GPT-6 Sol you could call functions through Chat Completions with
reasoning_effort: "none". GPT-6.1 Sol has nonone, and it supports Chat Completions only for requests without tools. Any Chat Completions plus tools path has to move to/v1/responses. - Drop sampling parameters. When effort isn't
none, OpenAI says to removetemperature,top_p, andtop_logprobs(pluslogprobsin Chat Completions). On GPT-6.1 Sol, that means every request.
Two smaller notes. If you change effort mid-conversation, use configuration_update items and keep the request-level reasoning.effort fixed so the cached prefix survives. And if you're jumping straight from GPT-5.5 or earlier, replace prompt_cache_retention with prompt_cache_options.ttl set to "30m".
curl: first request
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6.1-sol",
"reasoning": {"effort": "medium"},
"input": "Find the race condition in this Go worker pool and propose a minimal fix: ..."
}'
Python: a migration shim
This wraps old GPT-6 Sol request params so they're safe for GPT-6.1 Sol, then logs cache hits so you can watch the 5% rate do its job.
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from the environment
# GPT-6.1 Sol doesn't support these efforts; map them before switching.
EFFORT_MAP = {"none": "low", "minimal": "low"}
SAMPLING_KEYS = ("temperature", "top_p", "top_logprobs")
def to_gpt61_sol(params: dict) -> dict:
req = dict(params)
req["model"] = "gpt-6.1-sol"
effort = req.get("reasoning", {}).get("effort", "medium")
req["reasoning"] = {"effort": EFFORT_MAP.get(effort, effort)}
for key in SAMPLING_KEYS:
req.pop(key, None) # not allowed when effort is not "none"
return req
resp = client.responses.create(**to_gpt61_sol({
"input": "Summarize the failing tests and suggest the smallest fix.",
"reasoning": {"effort": "none"}, # legacy GPT-6 Sol setting
"temperature": 0.2,
}))
usage = resp.usage
cached = usage.input_tokens_details.cached_tokens
print(resp.output_text)
print(f"input={usage.input_tokens} cached={cached} output={usage.output_tokens}")
Node.js: compact before the 272K line
A cheap guard: estimate tokens and summarize the transcript before it crosses the surcharge threshold. The same code works against any OpenAI-compatible endpoint by changing baseURL. If you route through KissAPI or another gateway, confirm it exposes gpt-6.1-sol and the Responses endpoint before you move production traffic.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.API_KEY,
baseURL: process.env.API_BASE_URL, // e.g. https://api.openai.com/v1
});
// The surcharge starts above 272K input tokens; compact with headroom.
const COMPACT_AT = 250_000;
// Rough estimate: about 4 characters per token for English and code.
const estimateTokens = (text) => Math.ceil(text.length / 4);
export async function runTurn(history, userMsg) {
let context = history.join("\n\n");
if (estimateTokens(context) > COMPACT_AT) {
const summary = await client.responses.create({
model: "gpt-6.1-sol",
reasoning: { effort: "low" },
input: `Compress this agent transcript into the facts, decisions, and open TODOs needed to continue:\n\n${context}`,
});
context = summary.output_text;
}
return client.responses.create({
model: "gpt-6.1-sol",
reasoning: { effort: "medium" },
input: `${context}\n\n${userMsg}`,
});
}
Compaction rewrites your prefix, so the next turn pays one cache write. That's still far cheaper than paying 2x on every turn above 272K. For better estimates than four characters per token, paste real prompts into the token counter.
Which Model Should Get Your Traffic?
- Default agent and coding traffic: GPT-6.1 Sol at
medium. Raise effort tohighorxhighper task, not globally. - Your hardest tasks: keep a GPT-6 Astra escalation route, for example after two failed Sol attempts.
- Giant one-shot prompts above 272K: Claude Sonnet 5.5 bills its full 1M window at $2 / $10 with no surcharge, so it's cheaper there.
- High-volume simple work: GPT-6 Luna, at $0.10 input and $0.50 output per 1M tokens, is still the cheap lane.
My take: make GPT-6.1 Sol the default this week, but don't delete the Astra route. "Near-Astra" is an average across benchmarks, and the hardest 5% of your tasks is exactly where an average hides the gap. Pull 50 to 100 real tasks from your logs, run them on both models, and compare cost per passed task rather than cost per token. A model that needs a second attempt costs twice as much, whatever its rate card says.
FAQ
How much does the GPT-6.1 Sol API cost?
GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens. Cached input costs $0.10 per million tokens and cache writes cost $2.50. Batch and Flex are 50% off at $1 input and $5 output. Prompts above 272,000 input tokens are billed at 2x input and 1.5x output for the full request ($4 / $15).
What is the difference between GPT-6.1 Sol and GPT-6 Sol?
Both cost $2 input and $10 output per million tokens with a 1,050,000-token context window. GPT-6.1 Sol halves cached input to $0.10, has an April 30, 2026 knowledge cutoff, and scores closer to GPT-6 Astra on OpenAI's evals. It drops the none and minimal efforts, and its tool calling requires the Responses API.
Is GPT-6.1 Sol as good as GPT-6 Astra?
OpenAI says it's near-Astra for agentic coding, computer use, and professional work at one-fifth of Astra's per-token price ($2 / $10 vs $10 / $50). Those are OpenAI's own evals, and GPT-6 Astra is still OpenAI's most capable model. Test your hardest tasks on both before moving traffic.
Compare OpenAI and Claude Models With One Key
Create a free KissAPI account and A/B test models through one OpenAI-compatible endpoint before you commit production traffic.
Start Free