Skip to main content

Model comparison

Qwen3.8-2.4T-A95B vs Qwen3.8-27B

Two Qwen tiers compared on the figures that decide which one a workload actually needs.

Catalog record checked August 14, 2026; individual provider fields may change.

Qwen

Qwen3.8-2.4T-A95B

Frontier · Open weights

vs

Qwen

Qwen3.8-27B

Balanced · Open weights

AI model capability comparison
SpecificationQwen3.8-2.4T-A95BQwen3.8-27B
ProviderQwenQwen
TierFrontierBalanced
Context window262K262K
Max output131K131K
Input / 1M tokensNot verifiedNot verified
Output / 1M tokensNot verifiedNot verified
WeightsOpenOpen
Parameters2.4T total / 95B active (MoE)27B dense
Reasoning levelslow, medium, xhighlow, medium, xhigh
Modalitiestexttext, image, video
ReleasedAugust 12, 2026August 14, 2026

Prices are USD per million tokens at standard rates, excluding batch and caching discounts. Bold indicates the better figure where one is objectively better. Values we could not confirm from the provider are shown as “Not verified” rather than estimated.

FrontierOpen weights

Qwen3.8-2.4T-A95B

Official Hugging Face checkpoint for Qwen3.8 open weights — a text base model, not the hosted Qwen3.8-Max API.

Best for

  • Self-hosting the Qwen3.8 flagship
  • vLLM / SGLang serving
  • Evaluating the open Max-class weights

Watch out

Native context is 262,144 tokens (YaRN extends toward ~1,010,000). Hugging Face best-practice caps: 262,144 reasoning tokens and 131,072 final-response tokens. Thinking cannot be disabled. Qwen3.8-Max License: MaaS / AI Work Assistant businesses over $50M trailing 12-month revenue need a separate commercial licence. Self-hosting still needs datacentre-class memory. Do not treat this checkpoint as the multimodal 1M-context API.

BalancedOpen weights

Qwen3.8-27B

Apache 2.0 Qwen3.8 dense VLM for local and self-hosted work — native 262K context, image and video input, thinking on by default but can be turned off. Distinct from hosted Qwen3.8-Max.

Best for

  • Local multimodal agents
  • Self-hosting a dense 27B VLM
  • Apache 2.0 deployments

Watch out

The Hugging Face repo was created 2026-08-05; this row uses the 2026-08-14 card revision. Hosted 1M-context API is documented as coming soon. Vendor SWE-bench Pro 61.7 and DeepSWE 42.2 used a Claude Code harness, not swebench.com or the public Datacurve DeepSWE board. YaRN can extend context toward 1M.

Common questions

Qwen3.8-2.4T-A95B vs Qwen3.8-27B

Answered from the verified figures on this page rather than general guidance.

Which has the larger context window, Qwen3.8-2.4T-A95B or Qwen3.8-27B?

Both accept about 262K tokens of context, so document length will not decide between them.

Do Qwen3.8-2.4T-A95B and Qwen3.8-27B support the same reasoning levels?

Yes — both accept the same effort settings: "low", "medium", "xhigh". Higher effort costs more and takes longer, so start low and raise it only where output quality actually improves.

Should I use Qwen3.8-2.4T-A95B or Qwen3.8-27B?

Qwen3.8-2.4T-A95B is the frontier tier and Qwen3.8-27B the balanced tier. The useful question is whether your hardest task actually fails on the cheaper one — most production volume such as classification, extraction and summarisation does not.

Next step

Choosing between them

Tier and workload decide this more reliably than a leaderboard position does.

If both sit in the same tier, the decision usually comes down to context window and output price rather than headline capability — output tokens dominate real bills.

If one is a step up within the same provider, the useful question is whether your hardest task actually fails on the cheaper tier. Most production volume — classification, extraction, summarization — does not.

Choosing a harness rather than a model — Cursor, Copilot, Claude Code, Muse Code, or Lovable? Compare agentic harnesses · Latest releases.