Cost & pricing

GPT-6 Sol / Luna and Claude Opus 5.5: the September 22 price decks.

On one day, three providers republished the decks for the tier agent and coding workloads route to. OpenAI put GPT-6 Sol at $2.00 / $10.00 and GPT-6 Luna at $0.10 / $0.50 per million tokens; Anthropic put Claude Opus 5.5 at $4 / $20, below Claude Opus 5's $5 / $25; Google made Gemini 3.8 Flash TTS generally available.

Explore the live registry
GPT-6 Sol deck
$2 / $10

per 1M tokens · standard short context

Claude Opus 5.5
$4 / $20

−20% vs Claude Opus 5 at $5 / $25

Same-day releases
3 providers

OpenAI, Anthropic and Google · Sep 22, 2026

The launches at a glance

Prices per million tokens in USD on the provider’s own API, observed September 28, 2026. The Gemini TTS pair is text input / audio output; it is a different purchase unit from a chat model, not a cheaper chat model.

ModelProviderReleasedPrice (in / out)ContextMax outputDay-one surface
GPT-6 SolOpenAISep 22, 2026$2.00 / $10.001,050,000128,000OpenAI API — Responses and Chat Completions
GPT-6 LunaOpenAISep 22, 2026$0.10 / $0.501,050,000128,000OpenAI API — Responses and Chat Completions
Claude Opus 5.5AnthropicSep 22, 2026$4 / $201M128KClaude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS
Gemini 3.8 Flash TTSGoogleSep 22, 2026$0.50 / $9.00——Gemini API — Voices endpoint (text in / audio out)

Sources: developers.openai.com · API changelog (GPT-6 Sol and Luna release and rates) · platform.claude.com · Claude Opus 5.5 (Opus 5.5 price, context, output) · ai.google.dev · Gemini API changelog (Gemini TTS availability)

GPT-6 Sol is six published meters, not one rate

Standard short context is the deck a rate card records. The other tiers are published, billable and usually invisible in a normalized registry row.

MeterInputCached inputCache writesOutputNote
Standard · short context (≤272K)$2.00$0.20$2.50$10.00The deck a single per-deployment rate card records.
Standard · long context (>272K)$4.00$0.40$5.00$15.002x input and cache, 1.5x output, applied to the whole request.
Batch / Flex$1.00$0.10$1.25$5.0050% of Standard on both legs.
Fast mode$4.00$0.40$5.00$20.002x the applicable rate. Renamed from Priority Processing on Jul 30, 2026.

Sources: developers.openai.com · GPT-6 Sol model page (context, output, long-context rule, batch / flex / fast multipliers) · developers.openai.com · API pricing (deck rates) · developers.openai.com · API changelog (Luna at one twentieth of Sol)

The contract behind the deck

Where the model ID alone does not tell you what your integration will do.

Tool calling is a route choice, not a flag

GPT-6 Sol accepts text and image input and returns text, with reasoning effort from none to max. Built-in tools and function calling live on the Responses API; Chat Completions supports function calling only when reasoning_effort is set to none. A Chat Completions integration that wants tools at any reasoning level is a migration, not a model swap. Fine-tuning is not supported.

Residency is a processing-tier constraint here

Regional processing adds a 10% premium where it is available, and EU data residency for GPT-6 Sol and Luna is offered only with Standard processing — not with Batch, Flex or Fast mode. A long-context EU route is therefore a standard-tier purchase by construction.

An image bug was fixed three days later

On September 25, 2026 OpenAI shipped a fix for an image-encoding bug that degraded image understanding in GPT-6 Sol and GPT-6 Luna, and it recommends rerunning evaluations and retrying affected workflows where image inputs are used. If you benchmarked these models between September 22 and 25 on visual tasks, the number is stale.

Two models called “Sol”, two different decks

The registry has carried a model named gpt-5.6-sol for weeks. The new model is gpt-6-sol — a different ID, at a different price.

Model IDStandard inputStandard outputStatus of the deck
gpt-5.6-sol$4.00$20.00Promotional — “available at least through November 21, 2026”
gpt-6-sol$2.00$10.00Standard deck on the September 22, 2026 release

A cost model keyed on the word “Sol” is ambiguous

The new model’s standard input rate is half the promotional rate of the older one. Depending on which ID is actually called, a “Sol” line in a budget can be wrong by 2x in either direction — and the older ID’s deck is on a clock, so the two will not converge on their own.

Azure separates them by product family, not by name

Microsoft’s retail price list publishes the 5.6 meters under the product family Azure OpenAI GPT5 and the new ones under Azure OpenAI GPT6, with meter names 5.6-sol … and 6-sol …. The distinction is real in the meter data even where a deployment name is not.

Azure’s meter API priced the new Sol before its own price page

Two Microsoft surfaces currently disagree about GPT-6 Sol on Azure. The meters are the machine-readable source; the pricing page is the human one.

Microsoft’s pricing page still says “in processing”

The Azure OpenAI pricing page carries the notice that “GPT-6 Sol and Luna prices are currently in processing for publishing on this page”, and every GPT-6 Sol and Luna price cell on it is empty. The Azure Retail Prices API already returns the Azure OpenAI GPT6 meters for 6-sol in Standard, Data Zone and Priority Processing, in both short and long context.

Azure meter (Global)Azure rateOpenAI directVerdict
6-sol ShortCo Inp Std Gl$2.00$2.00Standard short-context input — matches direct
6-sol ShortCo Cd Inp Std Gl$0.20$0.20Cached input — matches direct
6-sol ShortCo Cd Wr Std Gl$2.50$2.50Cache writes — matches direct
6-sol ShortCo Opt Std Gl$10.00$10.00Standard short-context output — matches direct
6-sol LongCo Inp Std Gl$4.00$4.00Long-context input — matches direct
6-sol LongCo Opt Std Gl$15.00$15.00Long-context output — matches direct
6-sol ShortCo Inp PP Gl$4.00$4.00 (Fast mode)Azure Priority Processing = Fast mode
6-sol ShortCo Opt PP Gl$20.00$20.00 (Fast mode)Azure Priority Processing = Fast mode

Sources: prices.azure.com · Azure Retail Prices API · developers.openai.com · API pricing. Microsoft Foundry’s model catalog already lists gpt-6-sol and gpt-6-luna with a 2026-09-22 version date, and notes that the GPT-6 family requires a quota request on some subscription tiers: learn.microsoft.com · Foundry model catalog

Claude Opus 5.5: a 20% cut with four breaking changes

Anthropic moved the Opus line down a price tier on the same day, and attached request-shape changes that fail loudly.

The deck, and one rate that moved further than the headline

claude-opus-5-5 is priced at $4 / $20 per MTok against Claude Opus 5 at $5 / $25, with a 1M-token context window, 128K maximum output and adaptive thinking always on. Cache writes are $5 (5-minute) and $8 (1-hour) per MTok, and a cache read costs $0.20 — 5% of input rather than the 10% charged on most Claude models. For long agent loops with a stable system prompt, that cache line moves the effective rate more than the headline does.

Fast mode is cheaper than Opus 5’s, and first-party only

Fast mode (research preview) prices Claude Opus 5.5 at $8 / $40 per MTok, against $10 / $50 for Claude Opus 5 and Opus 4.8. It applies across the full context window including requests above 200k input tokens, and it is available on the Anthropic first-party API only — not on Claude Platform on AWS or partner-operated platforms, and not with the Batch API.

What returns an error on Opus 5.5

  • Thinking cannot be disabled: setting thinking to disabled or enabled returns a 400. Omit the field and steer depth with the effort parameter.
  • Forced tool use is rejected — tool_choice of any or tool returns a 400, as on Claude Fable 5.1. Use auto with strict tool use.
  • On the Claude API and Google Cloud, the earlier computer_20251124 tool is not accepted (use the newer toolset); on Amazon Bedrock the older tool keeps working — a real cross-platform divergence.
  • Text between tool calls comes back in thinking blocks whose text is empty at the default display setting, so a progress stream goes quiet between tool calls unless a display value is set.

Reach at day one, and a dated commitment

Claude Opus 5.5 shipped simultaneously on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry, so the release is reachable without leaving an existing cloud contract. Its listed retirement is not sooner than September 22, 2027.

Also official on September 22: Google’s audio surface

A new purchase unit, and a capacity control that is not a deprecation.

Gemini 3.8 Flash TTS and Flash-Lite TTS, with a Voices endpoint

Google made both TTS models generally available together with a Gemini API Voices endpoint, plus voice design, voice replication with consent verification and a library of 150+ voices. The flagship TTS model is priced at $0.50 per million text input tokens and $9.00 per million audio output tokens through December 31, 2026 — about $0.00225 per ten seconds of audio — and both rates double on January 1, 2027 ($1.00 / $18.00).

Gemini 2.5 access was restricted, but not deprecated

On September 18, 2026 Google limited access to the Gemini 2.5 models to users who had already used them, while stating that those models are not deprecated and will continue to be served, pointing new projects at 3.5 Flash-Lite or 3.8 Flash. A model can therefore be present, priced and undeprecated while no longer being provisionable for a new project — a distinction that a lifecycle-only view misses.

What changes for a routing or margin decision

  • Pin the model ID, not the family nickname. gpt-5.6-sol and gpt-6-sol are separate models at $4.00 and $2.00 standard input, and only the older one is on a promotional clock. A budget line labelled “Sol” describes neither precisely.
  • Treat a long-context request as a different purchase. Above 272K input tokens the whole request on GPT-6 Sol bills at 2x input and cache and 1.5x output — a single oversized prompt reprices the entire call, not just the excess. Chunk, summarize, or route to a smaller model before the threshold.
  • Check which Microsoft surface you are quoting. The Azure price page shows GPT-6 Sol and Luna as unpublished while the retail meter API already serves the rates. For a pre-signature Azure estimate, cite the meter API and note the page state rather than quoting either alone.
  • Re-baseline prompt-cache economics where Anthropic is in the mix. A 5% cache read on Claude Opus 5.5 is half the usual 10%, which changes the break-even point for long system prompts and agent-loop prefixes — and it stacks with the Batch API discount.
  • Validate request shape before wiring an Opus 5.5 fallback. Forced tool use and a disabled-thinking field both return 400 on this model, and the older computer-use tool is refused on the Claude API and Google Cloud but accepted on Bedrock. A fallback that works on one platform can fail on another.
  • Budget the January 1, 2027 step for Gemini TTS. Both the text and audio rates double on that date; a voice-agent cost model built on today’s audio rate understates 2027 by 2x.

This report covers a new model and a new deck. The earlier one covers one model’s meter spread, its promotional clock and the registry records that disagree with it.

Read GPT-5.6 Sol: one model, six meters, and a price that expires for the $5.00 / $30.00 pre-promotional deck, the registry’s alias records and the residency uplift on that model. This report covers the September 22, 2026 releases: two new GPT-6 models, an Opus price tier change and Google’s TTS pair.

Sources

All official provider pages observed September 28, 2026 (Europe/Paris). Azure meter rates read from the Azure Retail Prices API for the Azure OpenAI GPT6 product family on the same date.

Related live views: the OpenAI vs Azure AI deployment comparison and the GPT-5.6 Sol price-deck report render the current catalog records.

Deployment intelligence, explained

Why a same-day deck reset moves the routing default

Two things move independently when a provider ships a model: the model ID and the price deck attached to it. On September 22, 2026 both OpenAI and Anthropic republished the deck for the tier where coding and agent traffic lands — and OpenAI reused a name, “Sol”, that already belonged to a different model at a different price.

A published deck is a set of purchases, not a single rate. Standard, Batch, Flex, Fast mode, short and long context, and data residency are separate meters on the same model, so the headline rate and the rate a workload actually bills on can differ several-fold before residency is counted. Read the meter your traffic lands on, then the promotional clock attached to it.