Monthly roundup
August 2026 in AI models
A release-wave summary for readers who want the important changes without pretending every launch is a reason to switch.
The signal
12 releases plus 1 material status update
Z.ai, Qwen, OpenAI, DeepSeek, Google, Anthropic, SpaceXAI, Microsoft AI, Meta account for 12 tracked release events. 10 catalogued events link to a detailed comparison profile. 6 launch guides are available for deeper trade-off coverage. There is no single winner because the constraints are different.
At a glance
- Release events
- 12
- Profile links
- 10
What shipped and changed
- Release
GLM-5.3
Z.ai
Date note: Same base as GLM-5.2; gains are post-training. Direct API coming soon. Coding Plan access today (points quota; 5.2/5.1 requests route to 5.3). Open weights promised about two weeks after launch.
Z.ai GLM-5.3 announcement - Release
Qwen3.8-27B
Qwen
Date note: Apache 2.0 dense 27B VLM. Hugging Face repo created 2026-08-05; this event uses the 2026-08-14 card revision. Distinct from hosted Qwen3.8-Max.
Qwen3.8-27B on Hugging Face - Availability update
Daybreak Red and Daybreak Blue on Amazon Bedrock
OpenAI
Date note: Eligible-customer Bedrock access in US East (N. Virginia). Daybreak Red is GPT-5.6 Cyber; Daybreak Blue is GPT-5.6 Sol with defensive guardrails. Not a new catalog model.
AWS what's new - Release
DeepSeek Harness
DeepSeek
Date note: MIT-licensed v0.1 developer preview (`dsh`). Plugin-based agent runtime with a local web UI via `npx @deepseek-ai/dsh web`. Not a model profile.
DeepSeek Harness product page - General availability
DeepSeek-V4-Pro
DeepSeek
GA checkpoint rolled out on app, web, and API. The API id remains `deepseek-v4-pro`. Peak/off-peak pricing starts 2026-08-16 16:00 UTC.
DeepSeek API changelog - General availability
Gemini 3.7 Flash
Google
GA Flash workhorse for coding and agents. API id `gemini-3.7-flash`. Introductory paid rates $0.75 / $3.75 per million tokens through 2026-12-31, then $1.50 / $7.50.
Google Gemini 3.7 Flash announcement - Status update
Claude Sonnet 5 $2/$10 made standard
Anthropic
Date note: The previously scheduled 2026-09-01 step to $3/$15 will not occur.
Anthropic pricing docs - General availability
Grok 4.6
SpaceXAI
xAI's official model card records Grok 4.6 as the current code and chat default at $2/$6 with a 500K context window.
xAI Grok 4.6 model card - Availability update
MAI-Thinking-1
Microsoft AI
Date note: Public preview on Microsoft Foundry (256K context, 35B active / ~1T MoE). Closed weights. Not a generally available public API. Token list price not verified on Microsoft's pricing surface.
Microsoft AI announcement - Release
Qwen3.8-2.4T-A95B
Qwen
Date note: Open-weight checkpoint under the Qwen3.8-Max License. Distinct from the hosted Qwen3.8-Max API.
Qwen3.8-2.4T-A95B on Hugging Face - Release
Muse Glimmer 30B
Meta
Date note: Open-weight multimodal agentic distill of Muse Spark for ~24–32GB class local hardware; confirm the live HF/license card before redistribution.
Meta Muse Glimmer announcement - Release
Qwen3.8-Max
Qwen
Date note: Hosted Qwen3.8-Max API. The open-weight Qwen3.8-2.4T-A95B checkpoint is a separate 2026-08-12 event.
Alibaba Cloud announcement
Read the important launches
Gemini 3.7 Flash: the new coding and agent workhorse
Gemini 3.7 Flash arrives three weeks after 3.6 Flash as Google’s current Flash workhorse for coding and agents, with introductory API rates through the end of 2026.
Qwen
Qwen3.8-Max: hosted API versus the published open-weight checkpoint
Qwen3.8-Max is live as a hosted API with multimodal extras; the Qwen3.8-2.4T-A95B checkpoint is now on Hugging Face under the Qwen3.8-Max License and is a separate identity.
Qwen
Qwen3.8-27B: Apache 2.0 dense VLM for local work
Qwen3.8-27B is a 27B dense vision-language checkpoint under Apache 2.0 — a practical local path that is not the hosted Qwen3.8-Max API.
Meta
Muse Spark 1.2: a model co-trained with its coding harness
Muse Spark 1.2 is Meta's coding-focused model update, trained with Muse Code to improve long-horizon repository work and tool use.
Meta
Muse Code: Meta's beta terminal agent with persistent background workers
Muse Code is a macOS/Linux terminal agent built around persistent async workers, event-log replay, and a model co-trained for the harness.
Z.ai
GLM 5.3: Coding Plan flagship, API and weights still pending
GLM 5.3 is the current GLM Coding Plan default — same base as 5.2, with documented post-training gains. Direct token pricing and open weights are not published yet.