Z.ai
GLM 5.3 Flash
GLM 5.3 Flash is a Z.ai model with 1.05M context, $0.15 in / $0.50 out per million tokens, last verified September 3, 2026. Open weights: yes. Public-board scores are linked to the publisher; vendor-only figures stay in the watch-out, not on the board.
GLM 5.3 Flash is Z.ai's natively multimodal open-weight workhorse — 1M context, hybrid sparse/linear attention, near-flagship Intelligence Index at a fraction of the cost.
Catalog record checked September 3, 2026Individual provider fields may change
| Specification | GLM 5.3 Flash |
|---|---|
| Provider | Z.ai |
| Tier | Balanced |
| Context window | 1.05M |
| Max output | 131K |
| Input / 1M tokens | $0.15 |
| Output / 1M tokens | $0.50 |
| Weights | Open |
| Parameters | 320B total / 18B active (MoE) |
| Reasoning levels | Not verifiedUnverified |
| Modalities | text, image, video |
| License | MIT |
| Released | August 26, 2026 |
Pricing tiers: $0.15/$0.50 per MTok is the standard first-party and third-party rate (Z.ai docs; GMI, Novita, Together). A 50% promo tier runs $0.075/$0.25 — confirm which rate your account quotes.
Verified evidence
Published benchmark results
Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.
| Benchmark | Score | Measured | Source |
|---|---|---|---|
| Artificial Analysis Intelligence Index | 57 | 2026-08-26 | Artificial Analysis · View source |
| SWE-bench Verified | 92 | 2026-09-01 | vals.ai SWE-bench Verified leaderboard (independent, mini-swe-agent bash-only); not published by Z.ai · View source |
| Terminal-Bench 2.1 | 84.3 | 2026-08 | GLM-5.3-Flash HF model card eval results (vendor) · View source |
| DeepSWE 1.1 | 63.4 | 2026-08 | Z.ai GLM-5.3-Flash blog (vendor, mini-swe-agent, 400K context) · View source |
| Humanity's Last Exam | 55.3 | 2026-08 | GLM-5.3-Flash HF model card (vendor, with tools, full set) · View source |
Best for
- Cost-efficient long-context
- Multimodal input
- Coding agents
Watch out
Source receipts
Catalog figures for this model were checked against the following sources.
- Z.ai — GLM-5.3-Flash announcement (accessed 2026-08-29)
- getdeploying — GLM-5.3-Flash reference (accessed 2026-08-29)
- GMI Cloud — GLM-5.3-Flash analysis (accessed 2026-08-29)
Common questions
GLM 5.3 Flash
Answered from the verified figures on this page rather than general guidance.
How much does GLM 5.3 Flash cost per million tokens?
GLM 5.3 Flash is listed at $0.15 per million input tokens and $0.50 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate. $0.15/$0.50 per MTok is the standard first-party and third-party rate (Z.ai docs; GMI, Novita, Together). A 50% promo tier runs $0.075/$0.25 — confirm which rate your account quotes.
What is GLM 5.3 Flash's context window?
GLM 5.3 Flash accepts about 1.05M tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.
What is GLM 5.3 Flash best for?
GLM 5.3 Flash is a balanced tier from Z.ai. It suits cost-efficient long-context, multimodal input, coding agents. Self-hosting needs ~186GB GPU memory at 4-bit (multi-GPU). Third-party hosts charge more than Z.ai's own API.
Can I self-host GLM 5.3 Flash?
GLM 5.3 Flash publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.
Compare it
Featured head-to-head comparisons
A deliberately selected pairing for this model, rather than an automatically generated matrix.