Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
xAI: Grok 4.3 · Evaluations · Kaino
xAI: Grok 4.3 logo
model evaluation

xAI: Grok 4.3

xAI

Grok 4.3 is an xAI reasoning model with text and image inputs, text output, a 1,000,000-token context window, configurable reasoning, function calling, and structured outputs.

modellead-sourceopenrouter-modelsxaigrokreasoningmultimodallong-contextfunction-callingstructured-outputssource:docs.x.aiapi-model
72.4KAINO SCORERecommended
Evaluated Sep 7, 202611 reviews
Website Docs

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Pricing clarity88
  • Cost effectiveness82
  • Developer experience82
  • Technical capability80
  • Multimodal & I/O78
  • Risk & evidence70
  • Coding & agentic67
  • Speed & availability65
  • Reasoning & knowledge62
  • Adoption signal50

Kainotomic evaluation

Grok 4.3 has a strong documented API surface: text and image inputs, a 1M-token context window, configurable reasoning, function calling, and structured outputs. Its $1.25/M input, $0.20/M cached-input, and $2.50/M output pricing is materially more economical than premium frontier anchors such as Claude Opus 4.8, supporting a higher cost score. It is below Gemini 3.5 Flash on multimodal breadth because the supplied documentation only establishes image-to-text rather than broader native output modalities. Public performance evidence is mixed. LiveCodeBench’s mirrored Vals leaderboard reports 84.5% under high reasoning, but Terminal-Bench v2.1 reports 39.7% pass@1 and rank 26, while DeepSWE has no displayed Grok 4.3 result and no direct SWE-bench result was found. This supports capability above specialist multimodal Pegasus 1.5 but substantially below Claude Opus 4.8 and GPT-5.5 for coding agents. Artificial Analysis reports 128.4 output tokens/s, offset by a 22.38-second TTFT; availability/speed is therefore moderate rather than leading. The model’s reasoning score is held below Grok 3’s published anchor: the supplied Agent Arena snapshot places standard Grok 4.3 58th of 59 and High 52nd, despite the favorable LiveCodeBench result. Official docs and explicit regional pricing make integration and price terms clear. However, the model is new, public adoption evidence is limited, benchmark coverage is incomplete, and Artificial Analysis excerpts conflict between a current Intelligence Index of 38 and a launch-reported 53.

Strengths

  • 1M context with image input, function calling, structured outputs, and configurable reasoning
  • Low published token and cache pricing for a flagship reasoning API
  • High reported output throughput and a credible LiveCodeBench result

Caveats

  • Terminal-Bench v2.1 result is modest at 39.7% pass@1
  • Long 22.38-second TTFT weakens interactive latency
  • Agent Arena public-preference placement is near the bottom of the supplied snapshot