Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Grok 3 · Evaluations · Kaino
Grok 3 logo
model evaluation

Grok 3

xAI

xAI's Grok 3 model family for coding, technical Q&A, and developer workflows via the xAI API.

modelsource:x.aireasoning-modelcodingtechnical-qaxai
78.1KAINO SCORERecommended
Evaluated Jul 31, 202612 reviews
Website Docs

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Technical capability85
  • Coding & agentic82
  • Reasoning & knowledge82
  • Developer experience79
  • Multimodal & I/O76
  • Adoption signal76
  • Pricing clarity74
  • Speed & availability73
  • Cost effectiveness70
  • Risk & evidence70

Kainotomic evaluation

Grok 3 has strong but partly variant-specific evidence for coding and reasoning. xAI reports 79.4% on LiveCodeBench for Grok 3 (Think), while the independent LiveCodeBench listing places Grok-3-Mini (High) fourth at 71.3% Pass@1. This supports a score above GLM-4.6 on technical capability and reasoning, but below GPT-5.5 and Claude Opus 4.8, which have materially stronger calibrated agentic evidence. Artificial Analysis’ Intelligence Index of 18 and 1M-token context support solid general capability rather than frontier leadership. Coding evidence is credible but uneven across the family: Aider reports 49.3% for Grok 3 Mini Beta (high), and DeepSWE has no Grok 3 entry. Consequently, agentic-work scoring remains below Kimi K2.5 and DeepSeek-V3.2 despite the favorable LiveCodeBench result. Official documentation establishes API availability and developer relevance, but the supplied evidence does not substantiate broad tooling, reliability, or multimodal depth at the level of Gemini 3.5 Flash. At $4/M input and $20/M output tokens, Grok 3 is substantially less economical than low-cost coding models and well below DeepSeek-V3.2 on value, though less costly than premium frontier offerings. The cited fast-provider result (110.3 output tokens/s) is from SpaceXAI rather than a direct xAI availability measurement. Early 2025 Arena leadership and public visibility support adoption, but are dated. Vendor-reported benchmarks, family/variant conflation, absent DeepSWE coverage, and limited supplied safety evidence constrain confidence.

Strengths

  • Strong LiveCodeBench evidence for Grok 3 Think and Grok-3-Mini (High)
  • 1M-token context window
  • Documented xAI API model family with coding-oriented positioning
  • Meaningful early public-preference signal from Chatbot Arena coverage

Caveats

  • Benchmark results span Grok 3 Think, Grok-3-Mini, and Mini Beta variants rather than one uniform endpoint
  • No DeepSWE result is published for Grok 3
  • Aider performance is moderate rather than leading
  • Fast throughput evidence is for a third-party provider, not xAI directly