GPT-6.1 Sol Coding Agent API Guide (2026): Pricing, Context, and Production Routing

OpenAI released GPT-6.1 Sol for the API on September 29, 2026, shortly after DevDay. The pitch is straightforward: near-Astra capability for coding, computer use, and professional workflows, but at one-fifth of GPT-6 Astra's standard token price. That makes it less of a benchmark curiosity and more of a model you can actually put behind an agent.

The catch is that GPT-6.1 Sol rewards deliberate request design. Long prompts above 272,000 input tokens cost more, tool use belongs in the Responses API, and reasoning effort changes both latency and spend. Here's the practical setup I would use today.

Key takeaways
  • OpenAI released GPT-6.1 Sol on September 29, 2026, for coding, computer use, and professional API workloads.
  • GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens at standard rates for requests up to 272,000 input tokens.
  • GPT-6.1 Sol has a 1,050,000-token context window and a maximum output of 128,000 tokens.
  • GPT-6.1 Sol cached input costs $0.10 per million tokens, while input above 272,000 tokens uses a 2x input rate and a 1.5x output rate for the full request.

GPT-6.1 Sol API pricing and limits

The headline price is $2 per million input tokens and $10 per million output tokens. Cached input is $0.10 per million tokens, and cache writes are $2.50 per million tokens. Batch and Flex processing are priced at 50% below standard rates; Fast mode costs 2x standard. The detail that matters for coding agents is the long-context threshold: once a request exceeds 272,000 input tokens, OpenAI applies the higher rate to the full request rather than only the excess.

ModelInput / 1MOutput / 1MContext windowBest fit
GPT-6.1 Sol$2.00$10.001,050,000 tokensCoding agents and tools
GPT-6 Sol$2.00$10.001,050,000 tokensGeneral reasoning
GPT-6 Astra$10.00$50.001,050,000 tokensHardest frontier tasks

For a typical 20,000-token coding request that produces 2,000 output tokens, standard token cost is about $0.06 before any provider or gateway margin. A 300,000-token request is a different calculation: the long-context multiplier applies to the whole request, so trim stale history before you simply increase the context window.

First request: use the Responses API

GPT-6.1 Sol is designed around the Responses API rather than a legacy chat-completions-only integration. Start with a small, inspectable request and set reasoning effort explicitly so a production rollout isn't controlled by an invisible default.

curl https://api.openai.com/v1/responses \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6.1-sol",
    "reasoning": {"effort": "medium"},
    "input": "Review this diff for security bugs and return file:line findings."
  }'

For an OpenAI-compatible gateway, keep the same request shape and change the base URL. KissAPI is useful here when you want one billing and routing layer for multiple model families instead of rewriting your client every time you test an alternative.

Python: a safer coding-agent loop

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["KISSAPI_API_KEY"],
    base_url="https://api.kissapi.ai/v1"
)

def review_repo(summary: str) -> str:
    response = client.responses.create(
        model="gpt-6.1-sol",
        reasoning={"effort": "medium"},
        input=(
            "You are a security-minded code reviewer. "
            "Return only actionable findings with file and line references.

"
            + summary
        )
    )
    return response.output_text

print(review_repo("Diff summary: replaced JWT validation in auth/middleware.py"))

Keep the stable instructions short and put the current diff in the variable part of the request. Log input tokens, output tokens, latency, and the selected effort level. Without those four fields, it is hard to tell whether a slower result came from a larger repository, a more demanding reasoning setting, or a provider-side queue.

How GPT-6.1 Sol compares with the alternatives

OptionContextStandard priceBest forLimitation
GPT-6.1 Sol1.05M input context$2 / $10 per 1MAgentic coding with toolsAbove 272K input, full-request surcharge applies
GPT-6 Sol1.05M input context$2 / $10 per 1MStable general reasoningLess targeted than Sol for the newest coding workflow
GPT-6 Astra1.05M input context$10 / $50 per 1MMaximum capability5x input and output price

My default choice is Sol for multi-step coding, Astra only for tasks where a failed attempt costs more than the extra $40 per million output tokens, and Luna-class models for routing tests, extraction, and other bounded work. A model router should make that decision from task type and budget, not from a developer's favorite model name.

Production checklist

  1. Set a token budget. Cap output for review, classification, and patch-planning calls. A 128,000-token ceiling is a capability, not a sensible default.
  2. Trim history before 272K. Summarize old tool results and remove duplicate file contents before triggering long-context pricing.
  3. Use effort tiers. Start at medium. Escalate only when tests fail, the task touches security-sensitive code, or the agent needs another pass.
  4. Retry safely. Use idempotency keys for operations that can create tickets, commits, or deployments. Never blindly replay a tool call after a timeout.
  5. Keep a fallback. Route a second attempt to GPT-6 Sol or another approved model when the first call times out. Record the fallback in your usage data.

FAQ

What does GPT-6.1 Sol cost?

Standard pricing is $2 per million input tokens and $10 per million output tokens. Cached input is $0.10 per million tokens, and cache writes are $2.50 per million tokens.

Does GPT-6.1 Sol support long context?

Yes. The context window is 1,050,000 tokens, but requests above 272,000 input tokens receive higher full-request pricing. Treat one million tokens as a ceiling, not a target.

Should I use Chat Completions or Responses?

Use Responses for GPT-6.1 Sol when you need tools, reasoning controls, computer use, or multi-agent delegation. It also makes the agent's intermediate output easier to inspect.

Try GPT-6.1 Sol Without Rebuilding Your Stack

Create a free KissAPI account and use an OpenAI-compatible endpoint while you test model routing, token budgets, and fallbacks.

Start Free