Claude Sonnet 5.5 API Guide (2026): Pricing, Breaking Changes, and GPT-6 Sol Comparison

Anthropic released Claude Sonnet 5.5 on September 28, 2026. The pitch fits in one line: same price as Sonnet 5, 30%+ faster output, and up to 30% lower cost per task because the model needs fewer tokens to finish the same work.

What the launch post doesn't lead with is that Sonnet 5.5 ships with five breaking changes. If you swap claude-sonnet-5 for claude-sonnet-5-5 and deploy, some requests that worked yesterday will come back as 400s. This guide covers the price, the fixes, and how Sonnet 5.5 compares with Claude Opus 5.5 and GPT-6 Sol, which sells at the same list price.

Key takeaways

  • Anthropic released Claude Sonnet 5.5 (API model ID claude-sonnet-5-5) on September 28, 2026.
  • Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, with cache reads at $0.20 per million tokens, the same list price as Claude Sonnet 5.
  • Claude Sonnet 5.5 has a 1,000,000-token context window and 128,000 max output tokens, and the full 1M window is billed at standard rates with no long-context surcharge.
  • Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 (Claude Sonnet 5: 10.3%) and 55.5% on CursorBench 4.0 (Claude Sonnet 5: 34.1%), according to Anthropic.
  • On Claude Sonnet 5.5, thinking: {"type": "disabled"} and forced tool_choice of type any or tool return a 400 error; the replacements are between_tools thinking and tool_choice: auto with strict tools.

Claude Sonnet 5.5 Pricing and Limits

Nothing went up. Prices below are Anthropic's and OpenAI's published list prices as of September 29, 2026, in USD per 1M tokens.

ModelInputOutputCache readBatch (in / out)Context windowMax output
Claude Sonnet 5.5$2.00$10.00$0.20$1.00 / $5.001,000,000128,000
Claude Opus 5.5$4.00$20.00$0.20$2.00 / $10.001,000,000128,000
GPT-6 Sol$2.00$10.00$0.20$1.00 / $5.001,050,000 (922,000 max input)128,000
Claude Haiku 4.5$1.00$5.00$0.10$0.50 / $2.50200,00064,000

Two details matter more than the headline numbers. First, Sonnet 5.5 cache writes cost $2.50 per 1M tokens for the 5-minute cache and $4 for the 1-hour cache, and the minimum cacheable prompt dropped to 512 tokens (Sonnet 5 needed 1,024). Smaller system prompts can now hit the cache.

Second, long context. GPT-6 Sol charges 2x input and 1.5x output for the whole request once input passes 272K tokens. Claude Sonnet 5.5 bills the full 1M window at the standard rate. For a 600K-token input with 8K tokens of output:

Same sticker price, roughly half the bill. If you stuff whole repos or long contracts into one call, that gap adds up fast. The API cost calculator runs this math for your own traffic.

Sonnet 5.5 vs Opus 5.5 vs GPT-6 Sol

AttributeClaude Sonnet 5.5Claude Opus 5.5GPT-6 Sol
API model IDclaude-sonnet-5-5claude-opus-5-5gpt-6-sol
Price per 1M tokens (in / out)$2 / $10$4 / $20$2 / $10 (2x input, 1.5x output above 272K input)
Context window1,000,000 tokens1,000,000 tokens1,050,000 tokens
Default efforthighmediummedium
Thinking offbetween_tools (low, medium, high effort only)Not available (adaptive thinking always on)reasoning.effort: "none"
Terminal-Bench 4.070.6%66.4% (xhigh)Not publicly reported
CursorBench 4.055.5%57.8%Not publicly reported
Best forWell-scoped coding, bug fixes, docs, slides, spreadsheets, fast iterationOpen-ended work that needs sustained judgmentComplex coding and agent workflows on the OpenAI Responses API
Key limitationForced tool use and non-default sampling parameters return 400Twice the per-token price of Sonnet 5.5Chat Completions supports function calling only with reasoning_effort set to none

My take: Sonnet 5.5 should be the default for most coding agents now. Anthropic itself says Opus 5.5 is still stronger on complex, open-ended work, so keep Opus for tasks where Sonnet at high or xhigh still fails your evals. GPT-6 Sol fits stacks already built on the Responses API. These benchmarks are vendor-reported, so run your own.

The Five Breaking Changes (and the Fixes)

  1. thinking: {"type": "disabled"} returns 400. Adaptive thinking runs by default. The lowest setting is {"type": "between_tools"}, which skips up-front thinking. It only works at low, medium, or high effort, and it takes no other fields.
  2. Forced tool use is gone. tool_choice of type any or tool returns 400, including on the token counting endpoint. Use auto plus strict: true on the tool.
  3. Thinking blocks are tied to the model and the conversation. For accounts created on or after August 31, 2026, replaying a thinking block after editing earlier history returns 400. Keep conversation history append-only.
  4. computer_20251124 is rejected on the Claude API and Google Cloud. Move to computer_toolset_20260801.
  5. The advisor tool accepts fewer advisors. Claude Opus 4.8, Opus 4.7, and Sonnet 5 can no longer act as advisors for a Sonnet 5.5 executor.

One more change won't throw an error but will confuse users: longer notes the model writes between tool calls now come back as thinking blocks, which are empty at the default display setting. If your UI streams that text, it'll go quiet. Set display: "summarized" or use between_tools.

A clean first request (curl)

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-5-5",
    "max_tokens": 4000,
    "output_config": {"effort": "medium"},
    "messages": [
      {"role": "user", "content": "Find the off-by-one bug in this pagination helper: ..."}
    ]
  }'

Set effort explicitly. The API default is high, and thinking tokens bill as output. Anthropic recommends starting at medium for well-specified agentic coding and moving to high for harder tasks. The levels were recalibrated, so don't assume Sonnet 5's medium behaves like Sonnet 5.5's.

Python: turn off up-front thinking and read blocks by type

import os
from anthropic import Anthropic

client = Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])

resp = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=4000,
    thinking={"type": "between_tools"},  # replaces {"type": "disabled"}
    output_config={"effort": "medium"},  # between_tools only allows low/medium/high
    messages=[{"role": "user", "content": "Summarize this changelog in 5 bullets: ..."}],
)

# Don't use resp.content[0].text. The first block can be a thinking block.
text = "".join(block.text for block in resp.content if block.type == "text")
print(text)
print(resp.usage)

Python: replacing forced tool use

tools = [{
    "name": "record_ticket",
    "description": "Save a classified support ticket.",
    "strict": True,
    "input_schema": {
        "type": "object",
        "properties": {
            "priority": {"type": "string", "enum": ["low", "medium", "high"]},
            "summary": {"type": "string"}
        },
        "required": ["priority", "summary"],
        "additionalProperties": False  # required for strict tools
    }
}]

resp = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=2000,
    tools=tools,
    tool_choice={"type": "auto"},  # "any" and "tool" now return 400
    system="Always call record_ticket exactly once with your classification.",
    messages=[{"role": "user", "content": ticket_text}],
)

calls = [b for b in resp.content if b.type == "tool_use"]
if not calls:
    # auto means the model *can* skip the tool, so handle that path
    raise RuntimeError("No record_ticket call; retry or log for review")

With auto, the model is allowed to answer in plain text instead. The prompt has to say when to call the tool, and your code has to handle the case where it doesn't.

A/B Testing Sonnet 5.5 Against GPT-6 Sol

Since the two models cost the same per token, the real question is which one finishes your tasks in fewer tokens and fewer retries. The fastest way to find out is to send the same prompts to both through one OpenAI-compatible endpoint. With KissAPI, that's one key and one base URL for both Claude and GPT models:

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.KISSAPI_KEY,
  baseURL: "https://api.kissapi.ai/v1",
});

async function run(model, prompt) {
  const start = Date.now();
  const r = await client.chat.completions.create({
    model,
    messages: [{ role: "user", content: prompt }],
  });
  return { model, ms: Date.now() - start, usage: r.usage };
}

const prompt = "Write Vitest unit tests for parseDuration('1h30m').";
// Use the model IDs listed in your dashboard
for (const model of ["claude-sonnet-5-5", "gpt-6-sol"]) {
  console.log(await run(model, prompt));
}

Log usage per task, not per request. A model that takes three turns at the same price per token costs three times as much. Paste your typical prompts into the token counter first so your estimates start from real numbers.

Migration Checklist

If you only do two things today, set effort explicitly and grep your codebase for "disabled" and tool_choice. Those two account for most of the 400s you'll see.

FAQ

How much does the Claude Sonnet 5.5 API cost?

Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 per million tokens, 5-minute cache writes cost $2.50, and the Batch API is 50% off at $1 input and $5 output. The full 1,000,000-token context window is billed at standard rates.

Why does my Claude Sonnet 5 code return a 400 error on Claude Sonnet 5.5?

The usual causes are thinking: {"type": "disabled"} and forced tool_choice of type any or tool. Use between_tools to skip up-front thinking at low, medium, or high effort, and use tool_choice: auto with strict tools. On the Claude API and Google Cloud, computer_20251124 is also rejected in favor of computer_toolset_20260801.

Is Claude Sonnet 5.5 cheaper than GPT-6 Sol?

The list price is identical: $2 input and $10 output per million tokens. Above 272,000 input tokens, GPT-6 Sol bills 2x input and 1.5x output for the full request, while Claude Sonnet 5.5 keeps standard rates across its 1M window. For long-context requests, Sonnet 5.5 is cheaper.

Test Claude Sonnet 5.5 and GPT-6 Sol With One Key

Create a free KissAPI account and compare models through one OpenAI-compatible endpoint before you commit your production traffic.

Start Free