Moonshot AI says Kimi K3 is a 2.8-trillion-parameter, 1-million-context model aimed at long-horizon coding. The Terminal-Bench 2.1 tracker lists Kimi K3 at 88.3% task success, ahead of NVIDIA’s reported Nemotron 3 Ultra scores on the same benchmark.
Moonshot AI introduced Kimi K3 as a 2.8-trillion-parameter, 1-million-context model focused on long-horizon coding and agentic work.
Moonshot AI’s Kimi announcement says the company evaluated Kimi K3 on Terminal-Bench 2.1 using the KimiCode harness at maximum reasoning effort. The public Terminal-Bench 2.1 tracker on evals.report lists Kimi K3, from Moonshot AI, at 88.3% task success, with the result verified on July 17, 2026.
Terminal-Bench is designed to test whether AI systems can complete practical terminal-based tasks rather than only answer static questions. According to the Terminal-Bench 2.1 tracker, the benchmark covers agentic coding and system tasks in areas such as software engineering, administration, data processing, model training, and security, under controlled sandbox conditions with time and hardware limits.
That makes the Kimi K3 result notable because the score is attached to a benchmark that measures execution across multi-step computer-use tasks. It does not, by itself, establish broad superiority across all AI workloads, but it is a strong data point for terminal-based agentic performance.
The same evals.report Terminal-Bench 2.1 leaderboard lists NVIDIA Nemotron 3 Ultra at 56.4% task success. NVIDIA’s own NGC model page for NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 also reports a Terminal Bench 2.1 score of 56.4 in its agentic benchmark row.
NVIDIA Research’s technical report for Nemotron 3 Ultra reports Terminal Bench 2.1 scores of 56.4 for the BF16 version and 53.9 for the NVFP4 variant. Those figures broadly align with the leaderboard entry and indicate that NVIDIA’s published results are well below the 88.3% listed for Kimi K3 on this specific benchmark.
The comparison should be read carefully. Benchmarks can vary depending on harnesses, settings, tool access, inference configuration, and evaluation date. Moonshot AI’s Kimi K3 footnotes specify the KimiCode harness at maximum reasoning effort, while NVIDIA’s technical materials describe its Nemotron 3 Ultra variants and reported benchmark scores. The available sources support a direct comparison of the published Terminal-Bench 2.1 numbers, but not a universal ranking across all uses.
Moonshot AI’s published Kimi K3 materials position the model around long-context and long-horizon coding use cases, and the Terminal-Bench 2.1 score is consistent with that focus. A high result on terminal tasks suggests the model may be comparatively strong at executing structured, multi-step work in a command-line environment.
NVIDIA’s Nemotron 3 Ultra, meanwhile, is presented by NVIDIA Research as an open, efficient mixture-of-experts hybrid Mamba-Transformer model for agentic reasoning. NVIDIA’s documentation emphasizes model variants and efficiency considerations as well as benchmark results. Its lower Terminal-Bench 2.1 score does not erase those design goals, but it does show a meaningful gap on this particular task-success measurement.
For developers and AI infrastructure teams, the main takeaway is that agentic benchmarks are increasingly exposing differences that are not always visible in chat or general reasoning tests. Terminal-Bench 2.1 measures whether a system can complete tasks in a constrained computing environment, which is relevant for coding assistants, operations workflows, and automated data work.
Kimi K3’s 88.3% score, as listed by evals.report and tied to Moonshot AI’s own release notes, places it among the stronger reported performers on this benchmark. NVIDIA Nemotron 3 Ultra’s published 56.4% BF16 result provides a useful reference point, especially because it appears both in NVIDIA materials and on the same benchmark tracker.
As with any benchmark, the result is best treated as one source of evidence rather than a final verdict. Still, the published numbers show a clear performance difference on Terminal-Bench 2.1 and highlight how quickly agentic coding and terminal-use capabilities are becoming a competitive focus for open and open-weight AI systems.
Moonshot AI introduced Kimi K3 as a 2.8 trillion parameter, 1 million context model focused on long horizon coding and agentic work.
Kimi K3’s benchmark result Moonshot AI’s Kimi announcement says the company evaluated Kimi K3 on Terminal Bench 2.1 using the KimiCode harness at maximum reasoning effort.
The public Terminal Bench 2.1 tracker on evals.report lists Kimi K3, from Moonshot AI, at 88.3% task success, with the result verified on July 17, 2026.
Continue reading