Release watch

September 2026: the three-day frontier wave.

Anthropic, Google and OpenAI each shipped a new frontier-tier endpoint between September 1 and 3, 2026 — Claude Fable 5.1, Gemini 3.8 Flash and GPT-6 Astra. All three are absent from the current ALLM catalog snapshot, which predates the launches.

Explore the live registry
New releases
3

Sep 1–3, 2026 · Anthropic, Google, OpenAI

Top-of-line price
$10 / $50

per MTok · Fable 5.1 and GPT-6 Astra

Tracked in catalog
0 / 3

snapshot cat_2026_08_27 predates Sep 1

The three launches at a glance

Prices are input / output per million tokens on the provider’s own API, observed September 7, 2026. *Gemini 3.8 Flash intro pricing runs through December 31, 2026; from January 1, 2027 it becomes $1.50 / $7.50.

ModelProviderReleasedPrice (in / out)ContextMax outputALLM catalog
Claude Fable 5.1AnthropicSep 1, 2026$10 / $501M128KMissing — closest record: claude-fable-5
Gemini 3.8 FlashGoogleSep 2, 2026$0.75 / $3.75*1,048,57665,536Missing — closest record: gemini-3.7-flash
GPT-6 AstraOpenAISep 3, 2026$10 / $501,050,000128,000Missing — closest record: gpt-5.6-sol

Sources: platform.claude.com · Claude Fable 5.1 (Fable 5.1) · blog.google · Gemini 3.8 Flash (Gemini 3.8 Flash) · developers.openai.com · Changelog (GPT-6 Astra)

Per-model launch contract

Each launch changes a specific part of the operating contract. These are the constraints that alter an integration or a fallback — not benchmark claims.

Claude Fable 5.1 — Anthropic, Sep 1, 2026

  • Flagship tier at $10 / $50 per MTok with a 1M-token context window and 128K max output; adaptive thinking is always on with a default effort of high.
  • Prompt cache reads cost $0.25 per MTok (2.5% of input) — a quarter of the 10% rate on other Claude models — which changes the economics of long-system-prompt agent loops.
  • Available on day one on the Claude API, Amazon Bedrock (anthropic.claude-fable-5-1), Google Cloud, Microsoft Foundry and Claude Platform on AWS; retirement is committed “not sooner than September 1, 2027”.
  • Breaking changes versus Fable 5: forced tool use (tool_choice: any | tool) returns a 400 — use strict tool use or structured outputs — and thinking blocks are preserved only for the model that produced them or a newer one. The model also requires 30-day data retention and is not available under zero data retention.
  • Claude Mythos 5.1 ships the same capabilities, restricted to Project Glasswing participants.

Gemini 3.8 Flash — Google, Sep 2, 2026

  • The third Flash release in six weeks (3.6 on July 21, 3.7 on August 13, 3.8 on September 2), positioned as the “most intelligent workhorse model” with code execution, computer use (preview), search grounding, structured outputs and a 1,048,576-token input window.
  • Launches at the same introductory price as 3.7 Flash: $0.75 in / $3.75 out per MTok; the introductory price expires December 31, 2026, then $1.50 / $7.50 applies.
  • Thinking runs at low | medium | high; minimal is not supported and returns an error.
  • Gemini 3.7 Flash remains fully supported for efficiency-first workloads — 3.8 “works harder” and can spend more tokens at higher effort levels, so the choice is a cost/latency trade-off, not an automatic upgrade.
  • Gemini 3.8 Flash Cyber (vulnerability discovery and automated patching) is restricted to trusted defenders through the Fairwind Program.

GPT-6 Astra — OpenAI, Sep 3, 2026

  • OpenAI’s “most capable model” at $10 in / $50 out per MTok, with a 1,050,000-token context window, 128K max output and reasoning effort from low to max.
  • API surface changes: none reasoning effort is not supported, custom temperature / top_p and logprobs are not supported, and tool calling requires the Responses API — Chat Completions integrations must migrate before routing traffic.
  • Prompts above 272K input tokens are billed at 2x input and cache rates and 1.5x output for the full request; batch and flex are 50% of standard rates.
  • Microsoft Foundry lists it generally available with Standard and Provisioned Throughput in Global and US Data Zone geographies — same $10 / $50 short-context pricing globally, $11 / $55 in the US Data Zone, and a long-context tier at $20 / $75. The launch geography table lists no EU Data Zone.

The registry snapshot predates all three launches

ALLM catalog snapshot cat_2026_08_27 was observed August 26, 2026 — before any of these releases. The live registry tracks the prior generation on each surface (gpt-5.6-sol, gemini-3.7-flash, claude-fable-5 / claude-opus-5), so absence of a new model in the registry is a sync lag, not evidence that the endpoint does not exist.

What to do until the next catalog release

Plan against the official launch pages linked in the sources. The weekly deployment watch and the live comparison pages render the current snapshot and will not show these endpoints until the upstream catalog advances; do not read their absence as a shutdown or non-availability signal. Retry the registry after the next observed catalog release.

What changes for a routing decision

  • The cross-provider top tier is now price-symmetric at $10 / $50. Fable 5.1 and GPT-6 Astra share the same input/output price, so a provider fallback between Anthropic and OpenAI no longer trades price — it trades contract: Fable 5.1 forbids forced tool use and requires 30-day retention; Astra forbids temperature / top_p and needs the Responses API for tools. Validate both against your request shape before wiring a fallback.
  • Gemini Flash pricing has a January 1, 2027 cliff. 3.8 (and 3.7) Flash intro pricing doubles on that date ($0.75 → $1.50 in, $3.75 → $7.50 out). If a 2027 budget assumes today’s Flash rate, re-forecast; 3.7 Flash stays available for efficiency-first workloads if 3.8’s extra token spend does not fit your latency or cost envelope.
  • GPT-6 Astra long-context usage is expensive on every surface. OpenAI direct applies 2x input / 1.5x output above 272K input tokens; Foundry prices a separate long-context tier at $20 / $75 globally and $22 / $82.50 in the US Data Zone. Budget long-context agent workloads separately, and confirm EU Data Zone availability before committing EU traffic — the Foundry launch grid lists only Global and US geographies.
  • Same-week launches compress the migration window for the previous generation. The pattern to watch: when a flagship and a workhorse ship within 72 hours of each other across three providers, the deprecation notices for the prior tier (gpt-4o/o1 wave on October 1–21, claude-sonnet-4-5 “not sooner than” September 29) arrive faster than teams can test. Keep the existing Azure October wave report as the migration checklist while evaluating the new endpoints.

Sources

All official pages observed September 7, 2026 (Europe/Paris). ALLM catalog snapshot cat_2026_08_27 observed August 26, 2026.

Related live views: the October 2026 Azure retirement wave and the OpenAI vs Azure AI deployment comparison render the current catalog records.

Deployment intelligence, explained

Why same-week launches change the planning horizon

Between September 1 and 3, 2026, three frontier-tier releases landed: Claude Fable 5.1 on every Anthropic platform, GPT-6 Astra on the OpenAI API and Microsoft Foundry, and Gemini 3.8 Flash on the Gemini API. None of the three is in the current ALLM catalog snapshot (cat_2026_08_27, observed August 26, 2026), so an absence in the registry is a sync lag — not proof that an endpoint does not exist.

The useful comparison is the operating contract: price deck, context and output limits, effort and sampling constraints, tool-calling shape, platform reach at day one, and what stays supported behind each new release. The official launch pages linked in the sources are the planning source of record.