Monthly roundup
September 2026 in AI models
A release-wave summary for readers who want the important changes without pretending every launch is a reason to switch.
The signal
5 tracked releases across 5 providers
OpenAI, MBZUAI / IFM, Google, Meta, Anthropic account for 5 tracked release events. 4 catalogued events link to a detailed comparison profile. 4 launch guides are available for deeper trade-off coverage. There is no single winner because the constraints are different.
At a glance
- Release events
- 5
- Profile links
- 4
What shipped and changed
- Release
GPT-6 Astra
OpenAI
OpenAI launches GPT-6 Astra, its computer-use flagship and the first model gated at the Preparedness Framework's 'Critical' cybersecurity threshold — $10/$50 per MTok with a 1.05M-token context, 61.2 on the AA Intelligence Index and a self-reported 74.1% on DeepSWE 1.1.
OpenAI — Introducing GPT-6 Astra - ReleaseSecondary receipt
K2-Horizon Fleet (0.9B to 375B)
MBZUAI / IFM
MBZUAI releases fully open-source K2-Horizon fleet with MoVA attention, 512k context, and score 47 on Artificial Analysis Intelligence Index.
AI Release Tracker latest index (secondary source) - Release
Gemini 3.8 Flash + Flash Cyber
Google
Google unveils Gemini 3.8 Flash workhorse with 90.8% on Terminal-Bench 2.1 and companion defensive cyber edition via the Fairwind Program.
Google DeepMind — Gemini 3.8 Flash announcement - Release
Muse Spark 1.3
Meta
Meta Superintelligence Labs ships Muse Spark 1.3 with 75.4% on DeepSWE 1.1 and 1M token context for autonomous agent workflows.
Meta Superintelligence Labs — Muse Spark 1.3 announcement - Release
Claude Fable 5.1 + Mythos 5.1
Anthropic
Anthropic refreshes its flagship: Fable 5.1 debuts #1 on LiveBench (83.4, Max Effort) and the Artificial Analysis Intelligence Index (66) at unchanged $10/$50 pricing; Mythos 5.1 arrives alongside.
Anthropic pricing/models (LiveBench + Artificial Analysis cross-checked)
Read the important launches
OpenAI
GPT-6 Astra: OpenAI's computer-use flagship arrives gated
OpenAI's first GPT-6-family model is a computer-use flagship — and the first model it gates at the Preparedness Framework's 'Critical' cybersecurity threshold.
Anthropic
Claude Fable 5.1 and Mythos 5.1: the September refresh
Fable 5.1 lands at the top of the Artificial Analysis Intelligence Index at unchanged $10/$50 pricing — with a restricted-access Mythos 5.1 sibling for cyber and bio research.
Gemini 3.8 Flash: workhorse-class coding at a fifth of Opus cost
Google's third Flash release in six weeks is a coding and agentic workhorse — and it ties Claude Opus 5 at the top of DeepSWE 1.1 at roughly a fifth of the cost per task.
Meta
Muse Spark 1.3: Meta's self-reported DeepSWE leader
Meta's frontier refresh claims the highest DeepSWE 1.1 score yet published (75.4%) at $1.25/$4.25 — but the headline config is not broadly available yet.