Z.ai · Model release
GLM 5.3 Flash: the stealth model behind 'Ox Alpha' goes multimodal
Z.ai's first natively multimodal release shipped under an alias, then was claimed on 2026-08-26: 320B/18B MoE, 1M context, MIT licence, at a fraction of flagship pricing.
What changed
- The weights appeared anonymously as 'Ox Alpha' on OpenRouter/OpenCode about a week before Z.ai claimed them on 2026-08-26.
- Z.ai claims roughly 3× lower attention compute and a 4× smaller KV cache versus base GLM-5.3 at long context, from the hybrid sparse/linear-attention design.
- Pricing: Z.ai first-party API $0.08/$0.25 per million tokens; third-party hosts (GMI, Novita, Together) list $0.15/$0.50.
Best fit
- Cost-efficient long-context work
- Multimodal input pipelines
- Coding agents on a budget
What to be careful about
- Self-hosting needs roughly 186GB of GPU memory at 4-bit — multi-GPU only.
- The Intelligence Index 57 comes from Z.ai's own blog plus the Artificial Analysis board; treat vendor-adjacent numbers as provisional until independent replication.
- Text, image and video input only — no audio input and no built-in web search.
Dataset snapshot: GLM 5.3 Flash is also represented in the checked comparison dataset. Prices and benchmark figures carry their own field dates.