Skip to main content
AI Choice EngineAI Choice Engine
← All launch guides

Qwen · Model release

Qwen3.8-Flash-Next: an early look at the Qwen 4 architecture

Alibaba shipped an experimental preview of the architecture that will underpin Qwen 4: 125B total / 6B active MoE plus a 51B n-gram embedding table, at very low active-parameter cost.

Qwen

Released August 26, 2026

Compare it

What changed

  • The vendor-described n-gram component injects a 20-million-entry bigram/trigram table at an early layer — an unusual design aimed at cheap long-context decoding (not yet independently reviewed).
  • Native context is 262K, extensible to 1M; input accepts text, image, and video.
  • Weights are downloadable under the qwen-community-1.0 licence — more restrictive than the MIT/Apache norms on similar releases.

Best fit

  • Early architecture testing
  • High-throughput, low-active-parameter serving
  • Evaluating the Qwen 4 direction

What to be careful about

  • It is an experimental preview: expect breaking changes before the Qwen 4 family lands.
  • Benchmark claims (SWE-bench Pro 62.5, CoWorkBench 73.9) are vendor-reported with no independent replication yet.
  • QwenCloud pricing ($0.16/$0.47 reported) comes from a single secondary source — confirm before budgeting.
Dataset snapshot: Qwen 3.8 Flash Next is also represented in the checked comparison dataset. Prices and benchmark figures carry their own field dates.