Skip to main content

InclusionAI

Ling 3.0 Flash

InclusionAI's cost-focused Flash tier for high-frequency hybrid reasoning and agent loops at low list rates.

Catalog record checked August 13, 2026; individual provider fields may change.

AI model specification details
SpecificationLing 3.0 Flash
ProviderInclusionAI
TierBudget
Context window262K
Max outputNot verified
Input / 1M tokens$0.075
Output / 1M tokens$0.22
WeightsOpen
Parameters124B total / 5.1B active (MoE)
Modalitiestext
ReleasedJuly 24, 2026

Verified evidence

Published benchmark results

Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.

Artificial Analysis Intelligence Index

2026-08-08 · Artificial Analysis

Best for

  • High-volume agent loops
  • Cost-sensitive chat and extraction
  • Fast hybrid reasoning

Watch out

Native context is 256Ki tokens (262,144). Confirm serving provider and current list before committing volume.

Common questions

Ling 3.0 Flash

Answered from the verified figures on this page rather than general guidance.

How much does Ling 3.0 Flash cost per million tokens?

Ling 3.0 Flash is listed at $0.075 per million input tokens and $0.22 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate.

What is Ling 3.0 Flash's context window?

Ling 3.0 Flash accepts about 262K tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.

What is Ling 3.0 Flash best for?

Ling 3.0 Flash is a budget tier from InclusionAI. It suits high-volume agent loops, cost-sensitive chat and extraction, fast hybrid reasoning. Native context is 256Ki tokens (262,144). Confirm serving provider and current list before committing volume.

Can I self-host Ling 3.0 Flash?

Ling 3.0 Flash publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.

Compare it

Head-to-head model comparisons

These are the published pairings that put this model against a plausible alternative.