Skip to main content

Detailed model profiles

What has a profile in the last 28 days

This view shows releases with enough verified pricing, capability, and deployment data for a detailed comparison profile.

Window ends August 15, 2026. For announcements, previews, availability changes, and releases still awaiting a full profile, see the full release timeline. Need the comparison table too? Compare models · 2-min chooser · Stack cost calculator · Popular this week · Source receipts.

9 releases5 providers2 open-weight
Z.aiQwenGoogleSpaceXAIMetaDeepSeekAnthropicInclusionAI
  1. Qwen3.8-27B

    BalancedOpen weights

    Qwen

    Released yesterday

    Apache 2.0 Qwen3.8 dense VLM for local and self-hosted work — native 262K context, image and video input, thinking on by default but can be turned off. Distinct from hosted Qwen3.8-Max.

    • text
    • image
    • video

    Best for: Local multimodal agents · Self-hosting a dense 27B VLM · Apache 2.0 deployments

    Watch out: The Hugging Face repo was created 2026-08-05; this row uses the 2026-08-14 card revision. Hosted 1M-context API is documented as coming soon. Vendor SWE-bench Pro 61.7 and DeepSWE 42.2 used a Claude Code harness, not swebench.com or the public Datacurve DeepSWE board. YaRN can extend context toward 1M.

    Released
    August 14, 2026
    Input / 1M
    Not verified
    Output / 1M
    Not verified
    Context
    262K
    Max output
    131K
  2. Gemini 3.7 Flash

    BudgetClosed

    Google

    Released 2 days ago

    Google's 2026-08-13 Flash workhorse for coding and agents, three weeks after 3.6 Flash, at an introductory $0.75 / $3.75 per million tokens.

    • text
    • image
    • video
    • audio
    • pdf

    Best for: Coding agents · Multimodal agent loops · High-volume Google API work

    Watch out: Intro price doubles on 2027-01-01. DeepSWE 1.1 best-effort (high) is 65.3% at $2.18/task; medium scores 65.5% a few cents cheaper. 66K max output (65,536 tokens). Knowledge cutoff is March 2026 on some domains (January 2025 on others).

    Released
    August 13, 2026
    Input / 1M
    $0.75
    Output / 1M
    $3.75
    Context
    1.05M
    Max output
    66K
  3. Grok 4.6

    FrontierClosed

    SpaceXAI

    Released 3 days ago

    SpaceXAI's current code and chat default — same $2/$6 list as Grok 4.5, with a 500K context window and image input.

    • text
    • image

    Best for: Coding agents · Chat and knowledge work · Cost-sensitive frontier work

    Watch out: Token rates double above a 200K-token prompt. DeepSWE 1.1 best-effort for this family is the xhigh row at 66.7% (the catalog table rounds Pass@1 to 67%), not the higher-scoring medium config.

    Released
    August 12, 2026
    Input / 1M
    $2
    Output / 1M
    $6
    Context
    500K
    Max output
    Not verified
  4. Muse Glimmer 30B

    BalancedOpen weights

    Meta

    Released 5 days ago

    Meta's open-weight multimodal agentic model distilled from Muse Spark for local consumer hardware (~24–32GB class with 4-bit).

    • text
    • image

    Best for: Local agents · On-device coding + tool use · Privacy-sensitive multimodal work

    Watch out: Full BF16 needs far more than 24GB — plan on official GGUF / 4-bit packs (and mmproj for vision). Agentic quality ≠ frontier Muse Spark API; validate on your workflows.

    Released
    August 10, 2026
    Input / 1M
    $0
    Output / 1M
    $0
    Context
    131K
    Max output
    Not verified
  5. Muse Spark 1.2

    FrontierClosed

    Meta

    Released 10 days ago

    Meta Model API's current Muse Spark checkpoint (`muse-spark-1.2`). Meta documents it as the coding-focused Spark update used with Muse Code; maximum output is not published as a hard token cap on the current model page, so it stays unverified.

    • text
    • image
    • video
    • pdf

    Best for: Terminal coding agents · Long-horizon refactors · Meta Model API stacks

    Watch out: A cheaper contributor tier (`muse-spark-1.2-contributor`) permits training use of prompts and completions; standard-tier `muse-spark-1.2` does not. `reasoning_effort: none` is unsupported (HTTP 400). Weights are not marked open until an HF/license card is verified. For local agentic now, prefer Muse Glimmer 30B.

    Released
    August 5, 2026
    Input / 1M
    $1.25
    Output / 1M
    $4.25
    Context
    1.05M
    Max output
    Not verified
  6. Qwen3.8-Max

    FrontierClosed

    Qwen

    Released 12 days ago

    Alibaba's hosted Max-class API — 2.4T MoE with a 1M context window (max input 991,808), vision, and managed tools. Distinct from the Hugging Face Qwen3.8-2.4T-A95B checkpoint.

    • text
    • image
    • video

    Best for: Agentic coding · Multimodal knowledge work · Long-context delivery

    Watch out: This row is the hosted Qwen3.8-Max API, not the downloadable checkpoint. The open-weight Qwen3.8-2.4T-A95B card is catalogued separately and does not include the API's vision, non-thinking, or built-in tool extras.

    Released
    August 3, 2026
    Input / 1M
    $2
    Output / 1M
    $6
    Context
    1M
    Max output
    131K
  7. Claude Opus 5

    FrontierClosed

    Anthropic

    Released 22 days ago

    Anthropic's frontier tier, and ahead of Fable 5 on both recorded benchmarks at half the price.

    • text
    • image

    Best for: Agentic coding · Long-context analysis · Multi-step engineering work

    Watch out: Output tokens cost five times input — long generations dominate the bill.

    Released
    July 24, 2026
    Input / 1M
    $5
    Output / 1M
    $25
    Context
    1M
    Max output
    128K
  8. Gemini 3.6 Flash

    BudgetClosed

    Google

    Released 25 days ago

    Google's July 2026 Flash tier for multimodal volume work; 3.7 Flash is the newer coding/agent workhorse at the same introductory rate.

    • text
    • image
    • video
    • audio
    • pdf

    Best for: Multimodal pipelines · High-volume processing · Video and audio input

    Watch out: Introductory $0.75 input is 3.75× Luna; the $1.50 rate returns on 2027-01-01. 66K max output (65,536 tokens) is roughly half Luna's. Prefer 3.7 Flash for new coding-agent work unless you are pinned to 3.6.

    Released
    July 21, 2026
    Input / 1M
    $0.75
    Output / 1M
    $3.75
    Context
    1.05M
    Max output
    66K
  9. Gemini 3.5 Flash-Lite

    BudgetClosed

    Google

    Released 25 days ago

    Google's cheapest current Flash-class API tier for high-volume, latency-sensitive agent loops.

    • text
    • image
    • video
    • audio
    • pdf

    Best for: High-volume agents · Document pipelines · Cost-sensitive multimodal work

    Watch out: Cheaper than 3.6 and 3.7 Flash, not stronger on hard reasoning — pick it for throughput, not frontier coding.

    Released
    July 21, 2026
    Input / 1M
    $0.30
    Output / 1M
    $2.50
    Context
    1.05M
    Max output
    66K

Compare new releases

Head-to-head pages that include these models

Same-tier and same-provider pairings only — the comparisons a buyer would actually open.

Need a shortlist, not a changelog?

Release trackers tell you what shipped. Decision tools tell you what fits.

Use the AI coding harness finder or other guided tools when constraints (budget, latency, open weights) matter more than launch-day buzz.

Or jump straight to the full model comparison table.

How this list is built

Same verified dataset as the rest of the AI models section.

A model appears here only when its verified release date falls inside the rolling 28-day window. Missing dates are excluded on purpose — we do not invent launch days.

Pricing and benchmarks still carry their own verification dates on comparison pages.