Skip to main content

Detailed model profiles

What has a profile in the last 28 days

This view shows releases with enough verified pricing, capability, and deployment data for a detailed comparison profile.

Window ends August 15, 2026. For announcements, previews, availability changes, and releases still awaiting a full profile, see the full release timeline. Need the comparison table too? Compare models · 2-min chooser · Stack cost calculator · Popular this week · Source receipts.

5 releases4 providers5 open-weight
Z.aiQwenGoogleSpaceXAIMetaDeepSeekAnthropicInclusionAI
  1. Qwen3.8-27B

    BalancedOpen weights

    Qwen

    Released yesterday

    Apache 2.0 Qwen3.8 dense VLM for local and self-hosted work — native 262K context, image and video input, thinking on by default but can be turned off. Distinct from hosted Qwen3.8-Max.

    • text
    • image
    • video

    Best for: Local multimodal agents · Self-hosting a dense 27B VLM · Apache 2.0 deployments

    Watch out: The Hugging Face repo was created 2026-08-05; this row uses the 2026-08-14 card revision. Hosted 1M-context API is documented as coming soon. Vendor SWE-bench Pro 61.7 and DeepSWE 42.2 used a Claude Code harness, not swebench.com or the public Datacurve DeepSWE board. YaRN can extend context toward 1M.

    Released
    August 14, 2026
    Input / 1M
    Not verified
    Output / 1M
    Not verified
    Context
    262K
    Max output
    131K
  2. Qwen3.8-2.4T-A95B

    FrontierOpen weights

    Qwen

    Released 3 days ago

    Official Hugging Face checkpoint for Qwen3.8 open weights — a text base model, not the hosted Qwen3.8-Max API.

    • text

    Best for: Self-hosting the Qwen3.8 flagship · vLLM / SGLang serving · Evaluating the open Max-class weights

    Watch out: Native context is 262,144 tokens (YaRN extends toward ~1,010,000). Hugging Face best-practice caps: 262,144 reasoning tokens and 131,072 final-response tokens. Thinking cannot be disabled. Qwen3.8-Max License: MaaS / AI Work Assistant businesses over $50M trailing 12-month revenue need a separate commercial licence. Self-hosting still needs datacentre-class memory. Do not treat this checkpoint as the multimodal 1M-context API.

    Released
    August 12, 2026
    Input / 1M
    Not verified
    Output / 1M
    Not verified
    Context
    262K
    Max output
    131K
  3. Muse Glimmer 30B

    BalancedOpen weights

    Meta

    Released 5 days ago

    Meta's open-weight multimodal agentic model distilled from Muse Spark for local consumer hardware (~24–32GB class with 4-bit).

    • text
    • image

    Best for: Local agents · On-device coding + tool use · Privacy-sensitive multimodal work

    Watch out: Full BF16 needs far more than 24GB — plan on official GGUF / 4-bit packs (and mmproj for vision). Agentic quality ≠ frontier Muse Spark API; validate on your workflows.

    Released
    August 10, 2026
    Input / 1M
    $0
    Output / 1M
    $0
    Context
    131K
    Max output
    Not verified
  4. DeepSeek V4-Flash

    BudgetOpen weights

    DeepSeek

    Released 15 days ago

    Open-weight price-performance pick: MIT weights, 1M context, and API rates far below closed budget tiers.

    • text

    Best for: Cost-sensitive hosted agents · High-volume coding assist · Open-weight deployments with cluster VRAM

    Watch out: API list prices move to peak/off-peak from 2026-08-16 16:00 UTC (see pricing note). Self-hosting the ~284B MoE still needs roughly 90–170GB+ class VRAM depending on quant — not a laptop budget build. DeepSWE snapshot also uses the shared mini-swe-agent harness — verify cost, effort, and serving conditions before treating the result as a forecast.

    Released
    July 31, 2026
    Input / 1M
    $0.14
    Output / 1M
    $0.28
    Context
    1M
    Max output
    384K
  5. Ling 3.0 Flash

    BudgetOpen weights

    InclusionAI

    Released 22 days ago

    InclusionAI's cost-focused Flash tier for high-frequency hybrid reasoning and agent loops at low list rates.

    • text

    Best for: High-volume agent loops · Cost-sensitive chat and extraction · Fast hybrid reasoning

    Watch out: Native context is 256Ki tokens (262,144). Confirm serving provider and current list before committing volume.

    Released
    July 24, 2026
    Input / 1M
    $0.075
    Output / 1M
    $0.22
    Context
    262K
    Max output
    Not verified

Compare new releases

Head-to-head pages that include these models

Same-tier and same-provider pairings only — the comparisons a buyer would actually open.

Need a shortlist, not a changelog?

Release trackers tell you what shipped. Decision tools tell you what fits.

Use the AI coding harness finder or other guided tools when constraints (budget, latency, open weights) matter more than launch-day buzz.

Or jump straight to the full model comparison table.

How this list is built

Same verified dataset as the rest of the AI models section.

A model appears here only when its verified release date falls inside the rolling 28-day window. Missing dates are excluded on purpose — we do not invent launch days.

Pricing and benchmarks still carry their own verification dates on comparison pages.