Skip to main content
AI Choice EngineAI Choice Engine
← All launch guides

Z.ai · Model release

GLM 5.3 Flash: the stealth model behind 'Ox Alpha' goes multimodal

Z.ai's first natively multimodal release shipped under an alias, then was claimed on 2026-08-26: 320B/18B MoE, 1M context, MIT licence, at a fraction of flagship pricing.

Z.ai

Released August 26, 2026

Compare it

What changed

  • The weights appeared anonymously as 'Ox Alpha' on OpenRouter/OpenCode about a week before Z.ai claimed them on 2026-08-26.
  • Z.ai claims roughly 3× lower attention compute and a 4× smaller KV cache versus base GLM-5.3 at long context, from the hybrid sparse/linear-attention design.
  • Pricing: Z.ai first-party API $0.08/$0.25 per million tokens; third-party hosts (GMI, Novita, Together) list $0.15/$0.50.

Best fit

  • Cost-efficient long-context work
  • Multimodal input pipelines
  • Coding agents on a budget

What to be careful about

  • Self-hosting needs roughly 186GB of GPU memory at 4-bit — multi-GPU only.
  • The Intelligence Index 57 comes from Z.ai's own blog plus the Artificial Analysis board; treat vendor-adjacent numbers as provisional until independent replication.
  • Text, image and video input only — no audio input and no built-in web search.
Dataset snapshot: GLM 5.3 Flash is also represented in the checked comparison dataset. Prices and benchmark figures carry their own field dates.