Moonshot AI
Kimi K3
Open-weight frontier model scoring within a few points of the closed leaders on agentic coding.
Catalog record checked August 4, 2026; individual provider fields may change.
| Specification | Kimi K3 |
|---|---|
| Provider | Moonshot AI |
| Tier | Frontier |
| Context window | 1.05M |
| Max output | 1M |
| Input / 1M tokens | $3 |
| Output / 1M tokens | $15 |
| Weights | Open |
| Parameters | 2.8T total / 104B active (MoE) |
| Modalities | text, image, video |
| Released | July 16, 2026 |
Verified evidence
Published benchmark results
Each result keeps its source and measurement date visible. A missing benchmark is not treated as a zero.
Artificial Analysis Intelligence Index [max]
2026-08-08 · Artificial Analysis
Benchmark
DeepSWE 1.1 in context
The full local snapshot puts this model beside the wider field, including cost per completed task.
Local leader
Claude Opus 5 [max]
74%
Rows shown
24
Highest published reasoning effort per model (not best Pass@1)
Snapshot date
2026-08-13
Mirrored from deepswe.datacurve.ai
Better is toward the top-right (higher pass rate, lower cost). X-axis is reversed to match DeepSWE’s public chart. v1.1 uses average cost / tokens / steps; v1 uses published medians.
| # | Model | Pass@1 | Cost / task | Tokens / task | Steps / task |
|---|---|---|---|---|---|
| 1 | Claude Opus 5 [max] | 74% | $11.84 | 118k | 99 |
| 2 | GPT-5.6 Sol [max] | 73% | $8.39 | 60k | 61 |
| 3 | Claude Fable 5 [max] | 70% | $21.63 | 119k | 88 |
| 4 | GPT-5.6 Terra [max] | 70% | $4.95 | 72k | 76 |
| 5 | Kimi K3 [max] | 69% | $4.65 | 82k | 98 |
| 6 | GPT-5.6 Luna [max] | 67% | $3.03 | 73k | 102 |
| 7 | GPT-5.5 [xhigh] | 67% | $7.23 | 46k | 82 |
| 8 | Grok 4.6 [xhigh] | 67% | $5.50 | 71k | 87 |
| 9 | Gemini 3.7 Flash [high] | 65% | $2.18 | 107k | 125 |
| 10 | DeepSeek V4-Pro [max] | 63% | $0.24 | 106k | 155 |
| 11 | Claude Opus 4.8 [max] | 59% | $13.22 | 135k | 120 |
| 12 | Qwen3.8-Max [xhigh] | 58% | $3.73 | 95k | 111 |
| 13 | Muse Spark 1.2 [xhigh] | 55% | $3.70 | 99k | 101 |
| 14 | Claude Sonnet 5 [max] | 54% | $26.40 | 214k | 268 |
| 15 | Grok 4.5 [high] | 54% | $2.42 | 36k | 61 |
| 16 | DeepSeek V4-Flash [max] | 53% | $0.10 | 108k | 153 |
| 17 | Muse Spark 1.1 [xhigh] | 53% | $2.36 | 74k | 96 |
| 18 | GPT-5.4 [xhigh] | 52% | $5.65 | 71k | 70 |
| 19 | Gemini 3.6 Flash [high] | 47% | $4.42 | 96k | 117 |
| 20 | GLM 5.2 [max] | 44% | $3.92 | 78k | 129 |
| 21 | Gemini 3.5 Flash [high] | 36% | $3.45 | 76k | 105 |
| 22 | Kimi K2.7 Code | 31% | $2.82 | 59k | 149 |
| 23 | Claude Sonnet 4.6 [high] | 30% | $5.52 | 76k | 134 |
| 24 | Gemini 3.1 Pro [high] | 12% | $2.14 | 28k | 76 |
DeepSWE “Best” picks the highest published reasoning effort per model (not the highest pass rate). Small gaps may not be statistically meaningful — confirm on deepswe.datacurve.ai.
Best for
- Agentic coding without vendor lock-in
- Self-hosting at frontier quality
- Long-context work
Watch out
Common questions
Kimi K3
Answered from the verified figures on this page rather than general guidance.
How much does Kimi K3 cost per million tokens?
Kimi K3 is listed at $3 per million input tokens and $15 per million output tokens at standard rates. Output tokens usually dominate real bills, so weigh the output rate more heavily than the input rate.
What is Kimi K3's context window?
Kimi K3 accepts about 1.05M tokens of context. That only matters if you routinely send very long documents, large codebases, or multi-turn histories that approach that limit.
What is Kimi K3 best for?
Kimi K3 is a frontier tier from Moonshot AI. It suits agentic coding without vendor lock-in, self-hosting at frontier quality, long-context work. 2.8T parameters means self-hosting is a datacentre exercise, not a workstation one — open weights here mean provider choice, not local inference.
Can I self-host Kimi K3?
Kimi K3 publishes open weights, but self-hosting depends on the licence, hardware footprint, quantisation quality, and serving stack. A hosted API is often cheaper until you have measured throughput and concurrency on your own hardware.
Compare it
Head-to-head model comparisons
These are the published pairings that put this model against a plausible alternative.