Compare every Kimi model
All seven Kimi models in one table — context window, cache-hit and cache-miss input pricing, output pricing, capabilities and lifecycle status. Sort by any column, filter by the capability you need.
Last updated
| Kimi K4 | — | — | — | — | Unannounced |
|---|---|---|---|---|---|
| Kimi K3 | 1,048,576 | $0.30 | $3.00 | $15.00 | Available |
| Kimi K2.7 Code | 262,144 | $0.19 | $0.95 | $4.00 | Available |
| Kimi K2.7 Code HighSpeed | 262,144 | $0.38 | $1.90 | $8.00 | Available |
| Kimi K2.6 | 262,144 | $0.16 | $0.95 | $4.00 | Available |
| Kimi K2.5 | 262,144 | $0.10 | $0.60 | $3.00 | Closing |
| Kimi K2 | — | — | — | — | Discontinued |
A dash means Moonshot has not published that figure — we leave it blank rather than estimating.
FAQ
Choosing between Kimi models
Which Kimi model is best for coding?
Kimi K2.7 Code is the coding-specialised model, at $0.95 per million input tokens on a cache miss and $4.00 per million output tokens. Kimi K3 is stronger on the hardest problems and has a 1,048,576-token window, but its output tokens cost $15.00 per million — nearly four times as much.
Which Kimi model is cheapest?
Kimi K2.5 has the lowest published rates at $0.60 per million input tokens on a cache miss and $3.00 per million output tokens, but it is closed to new accounts and sunsets on August 31, 2026. Among models open to new users, Kimi K2.6 is cheapest on cache-hit input at $0.16 per million.
Is Kimi K2.7 Code HighSpeed worth double the price?
Only if latency is your binding constraint. HighSpeed runs the identical model at roughly 180 tokens per second, up to 260 in short contexts, and costs exactly twice as much on every tier. If your users are not waiting on the response in real time, the standard version gives identical output for half the money.
What is the difference between cache hit and cache miss pricing?
Cache-miss pricing applies the first time the model processes a stretch of input. Cache-hit pricing applies when a prompt prefix has been seen before and is served from Moonshot’s automatic context cache. The gap is large — on Kimi K3 it is $3.00 versus $0.30 per million tokens, a factor of ten.