Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Kimi K3 · Evaluations · Kaino
Kimi K3 logo
model evaluation

Kimi K3

Moonshot AI

Moonshot AI flagship 2.8T-parameter model with a 1M-token context window, native visual understanding, reasoning, long-horizon coding, and knowledge-work capabilities.

Moonshot AIkimi-k3
82.0KAINO SCORERecommended
Evaluated Jul 31, 202611 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Coding & agentic88
  • Pricing clarity85
  • Multimodal & I/O85
  • Developer experience84
  • Technical capability84
  • Reasoning & knowledge82
  • Speed & availability79
  • Cost effectiveness78
  • Adoption signal78
  • Risk & evidence77

Kainotomic evaluation

Kimi K3 has stronger public coding-agent evidence than the published Kimi K2.5 anchor: DeepSWE v1.1 reports 69%±5% with mini-swe-agent, fifth of 18, while Agent Arena places K3 Max third across 19,586 sessions. That supports an 88 coding-and-agentic score, above K2.5’s 86, but below GPT-5.5 and Claude Opus 4.8 because comparable broad software-engineering evidence is not supplied. Artificial Analysis’ Intelligence Index of 57 and Text Arena rank 11 (1486±10, preliminary) support solid but not frontier-leading general capability and reasoning. Official documentation supports a 1M-token context window, native vision, tool calling, JSON mode, structured output, caching, and reasoning-effort controls. This makes multimodal and API capability materially stronger than GLM-4.6 and DeepSeek-V3.2 in the published anchors, though the evidence does not establish parity with Opus 4.8’s broader multimodal score. First-party performance of roughly 35 tokens/s and 3.94s median first chunk is usable rather than leading; third-party routing can be much faster, but is not equivalent to first-party availability. At $3/M cache-miss input and $15/M output tokens, K3 is less economical than low-cost anchors such as DeepSeek-V3.2, though not priced like the most expensive frontier offerings. Pricing and API features are clearly documented. Evidence quality is moderately strong because official sources are supplemented by DeepSWE, Artificial Analysis, and Arena; however, no supplied SWE-bench, LiveCodeBench, Terminal-Bench, or Aider result substantiates further coding claims.

Strengths

  • DeepSWE v1.1 result of 69%±5%, fifth of 18 models
  • Third-place Agent Arena public outcome across 19,586 sessions
  • 1M-token context, native vision, and documented structured API controls
  • Clear published first-party token pricing

Caveats

  • Text Arena rank 11 is preliminary and preference-based rather than a capability benchmark
  • First-party throughput is moderate at about 35 tokens/s
  • Output pricing is relatively high at $15 per million tokens