Qwen
Qwen3.8-2.4T-A95B
Official Hugging Face checkpoint for Qwen3.8 open weights — a text base model, not the hosted Qwen3.8-Max API.
Catalog record checked August 14, 2026; individual provider fields may change.
| Specification | Qwen3.8-2.4T-A95B |
|---|---|
| Provider | Qwen |
| Tier | Frontier |
| Context window | 262K |
| Max output | 131K |
| Input / 1M tokens | Not verified |
| Output / 1M tokens | Not verified |
| Weights | Open |
| Parameters | 2.4T total / 95B active (MoE) |
| Modalities | text |
| Released | August 12, 2026 |
Best for
- Self-hosting the Qwen3.8 flagship
- vLLM / SGLang serving
- Evaluating the open Max-class weights
Watch out
Common questions
Qwen3.8-2.4T-A95B
Answered from the verified figures on this page rather than general guidance.
What is Qwen3.8-2.4T-A95B's context window?
Qwen3.8-2.4T-A95B accepts about 262K tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.
What is Qwen3.8-2.4T-A95B best for?
Qwen3.8-2.4T-A95B is a frontier tier from Qwen. It suits self-hosting the qwen3.8 flagship, vllm / sglang serving, evaluating the open max-class weights. Native context is 262,144 tokens (YaRN extends toward ~1,010,000). Hugging Face best-practice caps: 262,144 reasoning tokens and 131,072 final-response tokens. Thinking cannot be disabled. Qwen3.8-Max License: MaaS / AI Work Assistant businesses over $50M trailing 12-month revenue need a separate commercial licence. Self-hosting still needs datacentre-class memory. Do not treat this checkpoint as the multimodal 1M-context API.
Can I self-host Qwen3.8-2.4T-A95B?
Qwen3.8-2.4T-A95B publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.
Compare it
Head-to-head model comparisons
These are the published pairings that put this model against a plausible alternative.