Moonshot AI · stable
Moonshot AI's officially documented multimodal Kimi model; unverified K3 reports are excluded from this registry.
Specialist evidence
Source-native operational evidence for Kimi K2.5. Exact configurations remain separate. Costs are comparable only within the same selected benchmark/evaluation, and these rows never enter Overall Score.
| Evaluation | Exact configuration | Performance | Cost / task | Tokens / task | Execution |
|---|---|---|---|---|---|
| SWE-bench Owner Leaderboards · owner-current · multilingual SWE-bench Owner Leaderboards · checked 2026-08-29 | kimi-k2-5-thinking Kimi K2.5 | 67.3% | $0.693 | — | 50.4 calls |