Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Nemotron 3 Ultra · Evaluations · Kaino
Nemotron 3 Ultra logo
model evaluation

Nemotron 3 Ultra

NVIDIA

NVIDIA Nemotron 3 Ultra is an open 550B-parameter MoE language model with 55B active parameters, 1M context, configurable reasoning, and tool-use support.

nvidianemotron
74.6KAINO SCORERecommended
Evaluated Jul 31, 20269 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness85
  • Developer experience84
  • Speed & availability82
  • Technical capability80
  • Coding & agentic77
  • Risk & evidence75
  • Reasoning & knowledge74
  • Pricing clarity68
  • Adoption signal66
  • Multimodal & I/O55

Kainotomic evaluation

Nemotron 3 Ultra has unusually substantial official documentation: an open 550B/55B-active MoE checkpoint, text-only I/O, 1M-context claim, configurable reasoning, tool use, and NIM deployment/API materials. Its technical score is below GPT-5.5 and Claude Opus 4.8 because independent frontier-comparative evidence is limited, but above GLM-4.6 on architecture, context, and deployment scope. LiveCodeBench reporting lists Nemotron3U at 89% Pass@1; NVIDIA also reports SWE-Bench Verified in its technical report, though the supplied evidence does not provide a directly comparable score here. Coding and reasoning are therefore scored below DeepSeek-V3.2 and Kimi K2.5 despite the favorable LiveCodeBench result. In particular, Agent Arena places it #42 of 44, a material contrary signal for end-to-end agent workflows. Text-only I/O keeps multimodal capability aligned with DeepSeek-V3.2 and well below multimodal anchors. Artificial Analysis reports 203.8 output tokens/s and $0.675/$2.675 per million input/output tokens, supporting stronger speed and cost scores than premium closed models, subject to provider and configuration variation. NVIDIA's NIM documentation, compatible APIs, checkpoints, and NeMo repository make developer experience comparatively strong. Pricing clarity is weaker because cited AI Enterprise pricing is platform-level rather than a Nemotron 3 Ultra-specific production tariff. Public adoption remains modest versus OpenAI and Anthropic anchors. Evidence quality is reasonable but not top-tier: official claims are well documented, while no DeepSWE result was found, no Terminal-Bench/Aider result is supplied, and the preference-family source is Agent Arena rather than LMArena or Arena-Hard.

Strengths

  • Open checkpoint plus documented NIM deployment and OpenAI-/Anthropic-compatible API workflows
  • 1M-context claim, tool use, and configurable reasoning are documented by NVIDIA
  • Reported 89% LiveCodeBench Pass@1 and strong Artificial Analysis throughput/price figures

Caveats

  • Text-only model; no vision, audio, or other multimodal input is evidenced
  • Agent Arena rank (#42/44) conflicts with stronger coding claims
  • NVIDIA-reported SWE-Bench evidence is vendor-published and supplied material lacks a comparable numeric result