Kimi K3 has stronger public coding-agent evidence than the published Kimi K2.5 anchor: DeepSWE v1.1 reports 69%±5% with mini-swe-agent, fifth of 18, while Agent Arena places K3 Max third across 19,586 sessions. That supports an 88 coding-and-agentic score, above K2.5’s 86, but below GPT-5.5 and Claude Opus 4.8 because comparable broad software-engineering evidence is not supplied. Artificial Analysis’ Intelligence Index of 57 and Text Arena rank 11 (1486±10, preliminary) support solid but not frontier-leading general capability and reasoning. Official documentation supports a 1M-token context window, native vision, tool calling, JSON mode, structured output, caching, and reasoning-effort controls. This makes multimodal and API capability materially stronger than GLM-4.6 and DeepSeek-V3.2 in the published anchors, though the evidence does not establish parity with Opus 4.8’s broader multimodal score. First-party performance of roughly 35 tokens/s and 3.94s median first chunk is usable rather than leading; third-party routing can be much faster, but is not equivalent to first-party availability. At $3/M cache-miss input and $15/M output tokens, K3 is less economical than low-cost anchors such as DeepSeek-V3.2, though not priced like the most expensive frontier offerings. Pricing and API features are clearly documented. Evidence quality is moderately strong because official sources are supplemented by DeepSWE, Artificial Analysis, and Arena; however, no supplied SWE-bench, LiveCodeBench, Terminal-Bench, or Aider result substantiates further coding claims.