Cost & pricing

Claude Sonnet 5.5 and GPT-6.1 Sol: the September 28–29 workhorse reset.

Within two days, Anthropic and OpenAI both replaced the model tier that carries most coding and agent traffic. Claude Sonnet 5.5 shipped at $2 / $10 per million tokens with five breaking changes and put Claude Sonnet 4.5 on a November 30, 2026 retirement clock. GPT-6.1 Sol shipped at $2 / $10 with cache reads cut to $0.10 and a cache-write line that was not on the September 22 deck. A cross-provider fallback between the two no longer trades on price.

Explore the live registry
Claude Sonnet 5.5
$2 / $10

per 1M tokens · 1M context · effort high

GPT-6.1 Sol
$2 / $10

cache read halved to $0.10; new $2.50 cache-write line

Sonnet 4.5 retirement
Nov 30, 2026

announced Sep 30 · replacement Sonnet 5.5

Two workhorse releases, one day apart

USD per million tokens, standard short-context tier, observed October 5, 2026. Cache write is a separate purchase from a cache read; both providers now publish it.

ModelLabReleasedInputOutputCache readCache writeDefault effort
Claude Sonnet 5.5

claude-sonnet-5-5

AnthropicSep 28, 2026$2.00$10.00$0.20$2.50 (5m) · $4.00 (1h)high · adaptive
GPT-6.1 Sol

gpt-6.1-sol

OpenAISep 29, 2026$2.00$10.00$0.10$2.50reasoning settings

Sources: platform.claude.com · Claude Sonnet 5.5 and platform.claude.com · Release notes (Claude Sonnet 5.5 price, context, cache rates, effort) · developers.openai.com · API changelog and developers.openai.com · API pricing (GPT-6.1 Sol price and cache rates).

The GPT-6 deck, one week after the September 22 launch

Input / output per million tokens by processing tier, from OpenAI’s pricing page observed October 5, 2026. The third column is the model the September 22 note called GPT-6 Sol — the pricing page now sells the tier as gpt-6.1-sol.

Processing tiergpt-6-astragpt-6.1-solgpt-6-luna
Standard · short context (≤272K)$10.00 / $50.00$2.00 / $10.00$0.10 / $0.50
Standard · long context (>272K)$20.00 / $75.00$4.00 / $15.00$0.20 / $0.75
Batch / Flex$5.00 / $25.00$1.00 / $5.00$0.05 / $0.25
Fast mode$20.00 / $100.00$4.00 / $20.00$0.20 / $1.00
Ultrafast (new — Astra only)$60.00 / $300.00——

The same tier, a new ID, and a cheaper cache

The gpt-6-sol release note from September 22 priced cached input at $0.20. A week later the pricing page lists the tier as gpt-6.1-sol with cached input at $0.10 and a $2.50 cache-write line. Headline input and output are unchanged at $2.00 / $10.00, so a budget built on the headline is still right — but a prompt-cache model built on the old cache-read rate overstates the cached leg by 2×. The pattern repeats the September 22 finding: the ID a team pins is not guaranteed to be the ID the price page sells a week later.

Claude Sonnet 5.5: five breaking changes that fail loudly

Anthropic states that “code written for Claude Sonnet 5 can break on Claude Sonnet 5.5 in five ways.” Each one is a hard error or a shape change, not a warning.

The minimal change set

  • Up-front thinking is turned off with thinking {"type": "between_tools"} instead of "disabled", available at high effort or below.
  • Forced tool use returns a 400: tool_choice of any or tool is rejected, as on Fable 5.1 and Opus 5.5. Use auto with strict tool use.
  • Thinking blocks are tied to the model and the conversation.
  • On the Claude API and Google Cloud, the earlier computer_20251124 computer-use tool is not accepted.
  • The advisor tool rejects Claude Opus 4.8, Claude Opus 4.7, and Claude Sonnet 5 as advisors.

One more change that fails silently

Beyond the five, text emitted between tool calls now comes back inside thinking blocks whose text is empty at the default display setting. No request fails, but an application that streams that text to a user goes quiet between tool calls until it sets a display value that returns the text or turns off up-front thinking with between_tools. The model also rejects a non-default temperature, top_p or top_k with a 400, and the minimum cacheable prompt is 512 tokens.

Reach at day one, and a dated commitment

Claude Sonnet 5.5 shipped on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS — so the workhorse tier is reachable without leaving an existing cloud contract. Listed retirement is not sooner than September 28, 2027. On the Message Batches API it supports up to 300K output tokens under the output-300k-2026-03-24 beta header, with the usual 50% batch discount on input and output.

Claude Sonnet 4.5 is now on a two-month clock

The launch of the replacement opened a dated migration window for the outgoing workhorse.

Deprecated September 30, 2026 — retirement November 30, 2026

Anthropic notified developers using claude-sonnet-4-5-20250929 that it retires on the Claude API on November 30, 2026, with Claude Sonnet 5.5 as the recommended replacement. Microsoft’s Foundry retirement schedule already carries the same date for Azure-hosted claude-sonnet-4-5 (version 1) with the same replacement, so the two schedules corroborate. Separate from this: claude-haiku-4-5-20251001 is listed “not sooner than October 15, 2026” — a soft date, not a shutdown.

Also official on September 29, 2026

Two additions to the OpenAI surface that change an agent integration rather than a chat one.

Computer use reached the Agents API

Agents can now complete tasks in an OpenAI-hosted browser, with website access approvals and sign-in handled by your application. This is a capability addition to the managed agent harness, not a model change — an agent that browses inherits the same accountability boundary the harness already documented: the application owns approvals and credentials.

Ultrafast mode arrived for GPT-6 Astra — without EU residency

GPT-6 Astra gained a service_tier: "ultrafast" mode that reduces the time between generated output tokens. It is priced at $60.00 / $300.00 per MTok short context — 6× the standard rate — and is limited to global processing and US data residency; EU and other regional inference residency are not supported. It is the third latency tier on the same model (standard, Fast, Ultrafast), and the only one with a documented residency gap.

The September 25 fix: re-baseline image evaluations

On September 25, 2026 OpenAI shipped a fix for an image-encoding bug that degraded image understanding in GPT-6 Sol and GPT-6 Luna, and recommends rerunning evaluations and retrying affected workflows where image inputs are used. Any visual-task benchmark of those models measured between September 22 and 25 is stale, and a computer-use or vision feature validated in that window should be re-tested. This applies to the GPT-6 family, not to Claude Sonnet 5.5.

What changes for a routing or margin decision

  • A Sonnet-to-Sol fallback is now a contract swap, not a price swap. Both workhorse decks are $2.00 / $10.00 standard short context, so the deciding difference is what the API accepts: forced tool use and non-default sampling return 400 on Sonnet 5.5, while GPT-6.1 Sol routes tool calling through the Responses API. Validate both against your exact request shape before wiring the fallback.
  • Re-check pinning against the pricing page, not the launch note. The tier launched as gpt-6-sol is sold a week later as gpt-6.1-sol with a different cache-read rate. Reconcile the model IDs in your config against the provider’s current price page before assuming a pinned ID is still the cheapest correct route.
  • Cache writes are a budget line most cost models miss. Both providers now publish a separate cache-write rate ($2.50 on each). A prompt-cache strategy that assumed “cache reads are 10% of input” and ignored the write cost will understate spend on high-churn prefixes; a stable prefix that is read many times amortizes the write, a volatile one does not.
  • Start the Sonnet 4.5 migration now, not on November 30. Requests to a retired model fail. The replacement is a distinct contract with five breaking changes plus the between-tool-calls response-shape change, so the migration is code work, not an ID swap — and the two-month window closes before the end of the year.
  • Ultrafast is a latency purchase with a residency cost. At 6× standard input and no EU residency, it fits interruptible, latency-sensitive global/US workloads only. Keep residency-bound or cost-bound traffic on standard or Fast mode.

This report covers the following week’s workhorse releases and their contracts; the earlier one covers the September 22 launch decks and the registry records that disagree with them.

Read GPT-6 Sol / Luna and Claude Opus 5.5: the September 22 price decks for the flagship-tier decks and the gpt-5.6-sol vs gpt-6-sol naming collision. The ALLM catalog snapshot currently observed is cat_2026_09_22; it predates both Claude Sonnet 5.5 and GPT-6.1 Sol, so their absence from the live registry is a sync lag, not evidence that the endpoints do not exist.

Sources

All official provider pages observed October 5, 2026 (Europe/Paris).

Related live views: the September 22 price-deck report and the OpenAI vs Azure AI deployment comparison render the current catalog records.

Deployment intelligence, explained

Why a workhorse release changes more routing than a flagship

The flagship tier sets the ceiling; the workhorse tier carries the traffic. Claude Sonnet 5.5 and GPT-6.1 Sol both landed at $2.00 input / $10.00 output per million tokens in the last week of September 2026, so a cross-provider fallback between them no longer trades on price. It trades on the request contract: what an application may send, what the API rejects, and what a cache read costs.

A model ID is not a stable integration target on its own. OpenAI published the September 22 model as gpt-6-sol; the pricing deck one week later lists the tier as gpt-6.1-sol, at a halved cache-read rate and with a cache-write line that did not appear before. Read the meter your workload actually bills on, and the ID the provider currently sells — not the name from the launch note.