Model release
Qwen3.8-27B: Apache 2.0 dense VLM for local work
Qwen3.8-27B is a 27B dense vision-language checkpoint under Apache 2.0 — a practical local path that is not the hosted Qwen3.8-Max API.
Qwen
Read launch guideLaunch guides
Release dates tell you when something appeared. These guides explain what changed, who it fits, what remains unproven, and what to test before switching.
Model release
Qwen3.8-27B is a 27B dense vision-language checkpoint under Apache 2.0 — a practical local path that is not the hosted Qwen3.8-Max API.
Qwen
Read launch guideModel release
GLM 5.3 is the current GLM Coding Plan default — same base as 5.2, with documented post-training gains. Direct token pricing and open weights are not published yet.
Z.ai
Read launch guideModel release
Gemini 3.7 Flash arrives three weeks after 3.6 Flash as Google’s current Flash workhorse for coding and agents, with introductory API rates through the end of 2026.
Model release
Muse Spark 1.2 is Meta's coding-focused model update, trained with Muse Code to improve long-horizon repository work and tool use.
Meta
Read launch guideCoding harness
Muse Code is a macOS/Linux terminal agent built around persistent async workers, event-log replay, and a model co-trained for the harness.
Meta
Read launch guideModel release
Qwen3.8-Max is live as a hosted API with multimodal extras; the Qwen3.8-2.4T-A95B checkpoint is now on Hugging Face under the Qwen3.8-Max License and is a separate identity.
Qwen
Read launch guideModel release
The official V4-Flash build keeps the architecture but refreshes post-training for agents, coding, and tool use at unusually low API rates.
DeepSeek
Read launch guideModel release
Claude Opus 5 targets complex agentic coding, long-horizon work, and enterprise deliverables with a million-token context.
Anthropic
Read launch guideModel release
Gemini 3.6 Flash focuses on speed, multimodal input, and lower output usage for agentic and high-volume work.
Model release
Kimi K3 brings frontier ambitions, multimodal input, and open weights — but its scale changes what self-hosting means.
Moonshot AI
Read launch guideModel release
Grok 4.5 sits below the frontier labs on price with a large context window, image input, and optional live data tools — a mid-tier option when X or web grounding matters.
SpaceXAI
Read launch guideModel release
OpenAI's July family is a tier decision before it is a model decision: Sol for difficult work, Terra for balance, and Luna for volume.
OpenAI
Read launch guideModel release
Sonnet 5 looks cheaper per token than Opus 5, but long agentic runs can produce outsized output bills — the decision is completed-task cost, not the rate card alone.
Anthropic
Read launch guideModel release
GLM 5.2 pairs open weights with long context and mid-tier coding results — interesting for self-hosting teams that can carry integration and serving work.
Z.ai
Read launch guideModel release
Kimi K2.6 trades frontier context for open weights and low API rates — a sensible volume and self-hosting candidate when 256K context is enough.
Moonshot AI
Read launch guideModel release
Gemini 3.1 Pro targets very long documents and multimodal pipelines, but its limited-preview status and weak agentic-coding signal mean fit should be validated before a broad rollout.
Model release
Haiku 4.5 is Anthropic's budget route for classification and extraction, but its 200K context and 4.5-generation positioning mean it is a volume tool, not a frontier substitute.
Anthropic
Read launch guide