Skip to main content

Moonshot AI

Kimi K2.6

Open-weight mixture-of-experts model priced well below the closed frontier tiers.

Catalog record checked August 12, 2026; individual provider fields may change.

AI model specification details
SpecificationKimi K2.6
ProviderMoonshot AI
TierBudget
Context window262K
Max outputNot verified
Input / 1M tokens$0.95
Output / 1M tokens$4
WeightsOpen
Parameters1T total / 32B active (MoE)
Modalitiestext, image, video
ReleasedApril 21, 2026

Pricing tiers: Cache-hit input is $0.16 / MTok; cache-miss input is $0.95 / MTok; output $4.00 / MTok (Moonshot list).

Verified evidence

Published benchmark results

Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.

Artificial Analysis Intelligence Index

2026-08-08 · Artificial Analysis

Best for

  • Cost-sensitive volume
  • Self-hosting
  • Avoiding vendor lock-in

Watch out

262K context is among the smaller windows here — a real constraint on long-document work.

Common questions

Kimi K2.6

Answered from the verified figures on this page rather than general guidance.

How much does Kimi K2.6 cost per million tokens?

Kimi K2.6 is listed at $0.95 per million input tokens and $4 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate. Cache-hit input is $0.16 / MTok; cache-miss input is $0.95 / MTok; output $4.00 / MTok (Moonshot list).

What is Kimi K2.6's context window?

Kimi K2.6 accepts about 262K tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.

What is Kimi K2.6 best for?

Kimi K2.6 is a budget tier from Moonshot AI. It suits cost-sensitive volume, self-hosting, avoiding vendor lock-in. 262K context is among the smaller windows here — a real constraint on long-document work.

Can I self-host Kimi K2.6?

Kimi K2.6 publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.

Compare it

Head-to-head model comparisons

These are the published pairings that put this model against a plausible alternative.