Skip to main content
AI Choice EngineAI Choice Engine

Z.ai

GLM 5.3 Flash

GLM 5.3 Flash is a Z.ai model with 1.05M context, $0.15 in / $0.50 out per million tokens, last verified September 3, 2026. Open weights: yes. Public-board scores are linked to the publisher; vendor-only figures stay in the watch-out, not on the board.

GLM 5.3 Flash is Z.ai's natively multimodal open-weight workhorse — 1M context, hybrid sparse/linear attention, near-flagship Intelligence Index at a fraction of the cost.

Catalog record checked September 3, 2026Individual provider fields may change

AI model specification details
SpecificationGLM 5.3 Flash
ProviderZ.ai
TierBalanced
Context window1.05M
Max output131K
Input / 1M tokens$0.15
Output / 1M tokens$0.50
WeightsOpen
Parameters320B total / 18B active (MoE)
Reasoning levelsNot verifiedUnverified
Modalitiestext, image, video
LicenseMIT
ReleasedAugust 26, 2026

Pricing tiers: $0.15/$0.50 per MTok is the standard first-party and third-party rate (Z.ai docs; GMI, Novita, Together). A 50% promo tier runs $0.075/$0.25 — confirm which rate your account quotes.

Verified evidence

Published benchmark results

Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.

Published benchmark results for GLM 5.3 Flash
BenchmarkScoreMeasuredSource
Artificial Analysis Intelligence Index572026-08-26Artificial Analysis · View source
SWE-bench Verified922026-09-01vals.ai SWE-bench Verified leaderboard (independent, mini-swe-agent bash-only); not published by Z.ai · View source
Terminal-Bench 2.184.32026-08GLM-5.3-Flash HF model card eval results (vendor) · View source
DeepSWE 1.163.42026-08Z.ai GLM-5.3-Flash blog (vendor, mini-swe-agent, 400K context) · View source
Humanity's Last Exam55.32026-08GLM-5.3-Flash HF model card (vendor, with tools, full set) · View source

Best for

  • Cost-efficient long-context
  • Multimodal input
  • Coding agents

Watch out

Self-hosting needs ~186GB GPU memory at 4-bit (multi-GPU). Third-party hosts charge more than Z.ai's own API.

Source receipts

Catalog figures for this model were checked against the following sources.

Common questions

GLM 5.3 Flash

Answered from the verified figures on this page rather than general guidance.

How much does GLM 5.3 Flash cost per million tokens?

GLM 5.3 Flash is listed at $0.15 per million input tokens and $0.50 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate. $0.15/$0.50 per MTok is the standard first-party and third-party rate (Z.ai docs; GMI, Novita, Together). A 50% promo tier runs $0.075/$0.25 — confirm which rate your account quotes.

What is GLM 5.3 Flash's context window?

GLM 5.3 Flash accepts about 1.05M tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.

What is GLM 5.3 Flash best for?

GLM 5.3 Flash is a balanced tier from Z.ai. It suits cost-efficient long-context, multimodal input, coding agents. Self-hosting needs ~186GB GPU memory at 4-bit (multi-GPU). Third-party hosts charge more than Z.ai's own API.

Can I self-host GLM 5.3 Flash?

GLM 5.3 Flash publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.

Compare it

Featured head-to-head comparisons

A deliberately selected pairing for this model, rather than an automatically generated matrix.