Moonshot AI
Kimi K2.6
Open-weight mixture-of-experts model priced well below the closed frontier tiers.
Catalog record checked August 12, 2026; individual provider fields may change.
| Specification | Kimi K2.6 |
|---|---|
| Provider | Moonshot AI |
| Tier | Budget |
| Context window | 262K |
| Max output | Not verified |
| Input / 1M tokens | $0.95 |
| Output / 1M tokens | $4 |
| Weights | Open |
| Parameters | 1T total / 32B active (MoE) |
| Modalities | text, image, video |
| Released | April 21, 2026 |
Pricing tiers: Cache-hit input is $0.16 / MTok; cache-miss input is $0.95 / MTok; output $4.00 / MTok (Moonshot list).
Verified evidence
Published benchmark results
Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.
Artificial Analysis Intelligence Index
2026-08-08 · Artificial Analysis
Best for
- Cost-sensitive volume
- Self-hosting
- Avoiding vendor lock-in
Watch out
Common questions
Kimi K2.6
Answered from the verified figures on this page rather than general guidance.
How much does Kimi K2.6 cost per million tokens?
Kimi K2.6 is listed at $0.95 per million input tokens and $4 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate. Cache-hit input is $0.16 / MTok; cache-miss input is $0.95 / MTok; output $4.00 / MTok (Moonshot list).
What is Kimi K2.6's context window?
Kimi K2.6 accepts about 262K tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.
What is Kimi K2.6 best for?
Kimi K2.6 is a budget tier from Moonshot AI. It suits cost-sensitive volume, self-hosting, avoiding vendor lock-in. 262K context is among the smaller windows here — a real constraint on long-document work.
Can I self-host Kimi K2.6?
Kimi K2.6 publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.
Compare it
Head-to-head model comparisons
These are the published pairings that put this model against a plausible alternative.