Skip to main content

Detailed model profiles

What has a profile in the last 28 days

This view shows releases with enough verified pricing, capability, and deployment data for a detailed comparison profile.

Window ends August 15, 2026. For announcements, previews, availability changes, and releases still awaiting a full profile, see the full release timeline. Need the comparison table too? Compare models · 2-min chooser · Stack cost calculator · Popular this week · Source receipts.

5 releases3 providers2 open-weight
Z.aiQwenGoogleSpaceXAIMetaDeepSeekAnthropicInclusionAI
  1. Gemini 3.7 Flash

    BudgetClosed

    Google

    Released 2 days ago

    Google's 2026-08-13 Flash workhorse for coding and agents, three weeks after 3.6 Flash, at an introductory $0.75 / $3.75 per million tokens.

    • text
    • image
    • video
    • audio
    • pdf

    Best for: Coding agents · Multimodal agent loops · High-volume Google API work

    Watch out: Intro price doubles on 2027-01-01. DeepSWE 1.1 best-effort (high) is 65.3% at $2.18/task; medium scores 65.5% a few cents cheaper. 66K max output (65,536 tokens). Knowledge cutoff is March 2026 on some domains (January 2025 on others).

    Released
    August 13, 2026
    Input / 1M
    $0.75
    Output / 1M
    $3.75
    Context
    1.05M
    Max output
    66K
  2. DeepSeek V4-Flash

    BudgetOpen weights

    DeepSeek

    Released 15 days ago

    Open-weight price-performance pick: MIT weights, 1M context, and API rates far below closed budget tiers.

    • text

    Best for: Cost-sensitive hosted agents · High-volume coding assist · Open-weight deployments with cluster VRAM

    Watch out: API list prices move to peak/off-peak from 2026-08-16 16:00 UTC (see pricing note). Self-hosting the ~284B MoE still needs roughly 90–170GB+ class VRAM depending on quant — not a laptop budget build. DeepSWE snapshot also uses the shared mini-swe-agent harness — verify cost, effort, and serving conditions before treating the result as a forecast.

    Released
    July 31, 2026
    Input / 1M
    $0.14
    Output / 1M
    $0.28
    Context
    1M
    Max output
    384K
  3. Ling 3.0 Flash

    BudgetOpen weights

    InclusionAI

    Released 22 days ago

    InclusionAI's cost-focused Flash tier for high-frequency hybrid reasoning and agent loops at low list rates.

    • text

    Best for: High-volume agent loops · Cost-sensitive chat and extraction · Fast hybrid reasoning

    Watch out: Native context is 256Ki tokens (262,144). Confirm serving provider and current list before committing volume.

    Released
    July 24, 2026
    Input / 1M
    $0.075
    Output / 1M
    $0.22
    Context
    262K
    Max output
    Not verified
  4. Gemini 3.6 Flash

    BudgetClosed

    Google

    Released 25 days ago

    Google's July 2026 Flash tier for multimodal volume work; 3.7 Flash is the newer coding/agent workhorse at the same introductory rate.

    • text
    • image
    • video
    • audio
    • pdf

    Best for: Multimodal pipelines · High-volume processing · Video and audio input

    Watch out: Introductory $0.75 input is 3.75× Luna; the $1.50 rate returns on 2027-01-01. 66K max output (65,536 tokens) is roughly half Luna's. Prefer 3.7 Flash for new coding-agent work unless you are pinned to 3.6.

    Released
    July 21, 2026
    Input / 1M
    $0.75
    Output / 1M
    $3.75
    Context
    1.05M
    Max output
    66K
  5. Gemini 3.5 Flash-Lite

    BudgetClosed

    Google

    Released 25 days ago

    Google's cheapest current Flash-class API tier for high-volume, latency-sensitive agent loops.

    • text
    • image
    • video
    • audio
    • pdf

    Best for: High-volume agents · Document pipelines · Cost-sensitive multimodal work

    Watch out: Cheaper than 3.6 and 3.7 Flash, not stronger on hard reasoning — pick it for throughput, not frontier coding.

    Released
    July 21, 2026
    Input / 1M
    $0.30
    Output / 1M
    $2.50
    Context
    1.05M
    Max output
    66K

Compare new releases

Head-to-head pages that include these models

Same-tier and same-provider pairings only — the comparisons a buyer would actually open.

Need a shortlist, not a changelog?

Release trackers tell you what shipped. Decision tools tell you what fits.

Use the AI coding harness finder or other guided tools when constraints (budget, latency, open weights) matter more than launch-day buzz.

Or jump straight to the full model comparison table.

How this list is built

Same verified dataset as the rest of the AI models section.

A model appears here only when its verified release date falls inside the rolling 28-day window. Missing dates are excluded on purpose — we do not invent launch days.

Pricing and benchmarks still carry their own verification dates on comparison pages.