Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Kimi K2.7 Code · Evaluations · Kaino
Kimi K2.7 Code logo
model evaluation

Kimi K2.7 Code

Moonshot AI

Open-weight, coding-focused Kimi model for long-horizon software engineering, tool use, large-context codebase work, and coding agents.

Moonshot AIKimimultimodal
80.5KAINO SCORERecommended
Evaluated Jul 31, 202615 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Pricing clarity88
  • Cost effectiveness85
  • Developer experience84
  • Multimodal & I/O82
  • Technical capability82
  • Coding & agentic82
  • Speed & availability80
  • Risk & evidence77
  • Reasoning & knowledge75
  • Adoption signal70

Kainotomic evaluation

Kimi K2.7 Code has credible coding capacity, but public evidence is mixed across task types. NVIDIA’s model card reports 74.1 SWE-bench Verified for the native INT4 baseline, while DeepSWE reports 31%±1% pass@1 and 16th of 18 displayed entries. Its Long-Horizon Terminal-Bench result—11th, 0.367 mean reward—also supports a score below the strongest coding anchors. It therefore trails Claude Opus 4.8 (95 coding) and the prior Kimi K2.5 evaluation (86), while materially exceeding Gemma 3’s less agent-oriented 42. Official materials substantiate a 262,144-token context, 1T-parameter/32B-active MoE design, tool calling, JSON/partial modes, OpenAI-compatible API, and image/video input. This makes multimodal I/O comparable to K2.5 but below Gemini 3.5 Flash’s stronger 92. Artificial Analysis lists 47.5 output tokens/s; availability and speed are solid rather than category-leading. Published $0.19 cached-input, $0.95 uncached-input, and $4 output per million tokens are unusually competitive against premium frontier models, with clear official pricing. Public preference is modest: Code Arena places it #27 with 4,677 votes and Agent Arena #22 of 38. LiveCodeBench was checked but has no listed result; the model card’s proprietary coding benchmarks are useful context, not independent confirmation. Open weights, Hugging Face materials, and API compatibility support developer experience, but the divergent SWE-bench, DeepSWE, and terminal results warrant a conservative evidence score.

Strengths

  • Officially documented 256K context, tool calling, OpenAI-compatible API, and image/video input
  • Strong reported SWE-bench Verified result from NVIDIA’s model card
  • Competitive published token pricing, especially cached input
  • Open-weight distribution and deployment materials

Caveats

  • DeepSWE performance is weak relative to the displayed field
  • Long-Horizon Terminal-Bench ranking and reward are mid-pack
  • No official LiveCodeBench listing was found
  • Arena rankings indicate limited public-preference strength versus leading coding models