Meta
Llama 4 Scout
Llama 4 Scout is a Meta model with 10M context, Not verified in / Not verified out per million tokens, last verified September 4, 2026. Open weights: yes. Public-board scores are linked to the publisher; vendor-only figures stay in the watch-out, not on the board.
Llama 4 Scout is Meta's open-weight 109B/17B MoE with a 10M-token extended context (maintenance-mode — Meta ended Llama development for Muse).
Catalog record checked September 4, 2026Individual provider fields may change
| Specification | Llama 4 Scout |
|---|---|
| Provider | Meta |
| Tier | Budget |
| Context window | 10M |
| Max output | 131K |
| Input / 1M tokens | Not verifiedUnverified |
| Output / 1M tokens | Not verifiedUnverified |
| Weights | Open |
| Parameters | 109B total / 17B active (MoE) |
| Reasoning levels | Not verifiedUnverified |
| Modalities | text, image |
| License | Llama 4 Community License |
| API model id | llama-4-scout |
| Released | April 5, 2025 |
Pricing tiers: Open-weight (Llama 4 Community License) — self-host free; 10M-token context via iRoPE (256K pretrained, length-generalised). No first-party Meta API: the Llama API Public Preview was retired 2026-07-06; Meta's endpoint serves Muse models only.
Verified evidence
Published benchmark results
Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.
| Benchmark | Score | Measured | Source |
|---|---|---|---|
| Artificial Analysis Intelligence Index | 40 | 2026-08-14 | Artificial Analysis · View source |
Best for
- Very long-context retrieval
- Open-weight self-hosting
- Single-GPU (int4)
Watch out
Source receipts
Catalog figures for this model were checked against the following sources.
- Meta — Llama 4 (accessed 2026-08-29)
- Promptfoo — Llama API provider docs (Public Preview retired 2026-07-06) (accessed 2026-09-04)
- Groq — deprecations (llama-4-scout-17b-16e-instruct EOL 2026-07-17) (accessed 2026-09-04)
Common questions
Llama 4 Scout
Answered from the verified figures on this page rather than general guidance.
What is Llama 4 Scout's context window?
Llama 4 Scout accepts about 10M tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.
What is Llama 4 Scout best for?
Llama 4 Scout is a budget tier from Meta. It suits very long-context retrieval, open-weight self-hosting, single-gpu (int4). 10M context is extended capability; practical use starts at 256K pretrained window. Meta ended Llama development for the Muse family, and Groq retired its Scout endpoint on 2026-07-17 — treat Llama 4 as maintenance-mode open weights.
Can I self-host Llama 4 Scout?
Llama 4 Scout publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.