Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
GPT-5.4 · Evaluations · Kaino
GPT-5.4 logo
model evaluation

GPT-5.4

OpenAI

OpenAI frontier model for professional reasoning, coding, tool use, and agentic workflows across ChatGPT, API, and Codex.

modelOpenAIChatGPT
81.2KAINO SCORERecommended
Evaluated Jul 31, 202612 reviews
Website Docs

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Developer experience89
  • Coding & agentic89
  • Pricing clarity88
  • Technical capability88
  • Multimodal & I/O82
  • Reasoning & knowledge81
  • Adoption signal80
  • Risk & evidence78
  • Cost effectiveness69
  • Speed & availability68

Kainotomic evaluation

GPT-5.4 has strong public evidence for coding and agent use, though it is not a clear category leader. Terminal-Bench reports 77.3% with Codex CLI on Terminal-Bench 2.1, and a third-party LiveCodeBench reporting page lists Codex GPT-5.4 at 91.8% pass@1. Its 1.1M-token context and Artificial Analysis Intelligence Index of 51 support a high technical score. It remains below the published GPT-5.5 and Claude Opus 4.8 anchors in technical and agentic work, whose evaluations carry stronger top-end evidence. The coding score is moderated by DeepSWE: gpt-5.4 xhigh scored 52% ±2% and ranked 12th of 18 displayed configurations. Current Text Arena placement (#12 overall, #8 Expert, #12 Hard Prompts) supports solid but not leading reasoning/public-preference performance; this is stronger than lower-reasoning anchors such as GPT-5.6 Terra, but below GPT-5.5, GPT-5.6 Sol, and Opus 4.8. Official documentation supports API, tool-use, and safety-material availability, while the supplied evidence does not substantiate detailed modality behavior beyond official capability claims. At $2.50/M input and $15/M output, GPT-5.4 is materially less cost-effective than DeepSeek-V3.2 or Gemma 3, despite the large context window. Reasoning-mode latency is a constraint: Artificial Analysis reports 121.68s TTFT, versus 0.80s for non-reasoning; output throughput is high after generation begins. Pricing documentation, model docs, system-card material, and Codex integration make the developer experience strong. Evidence quality is good but limited by missing directly evidenced SWE-bench results and the discrepancy between official LiveCodeBench absence and a third-party Codex listing.

Strengths

  • 77.3% on Terminal-Bench 2.1 with Codex CLI
  • 1.1M-token context window
  • Strong API, tool-use, Codex, pricing, and system-card documentation
  • High reported post-start output throughput of 143.7 tokens/s in xhigh mode

Caveats

  • DeepSWE result is 52% ±2%, ranking 12th of 18 displayed configurations
  • Reasoning xhigh TTFT is 121.68 seconds
  • $15/M output pricing weakens value against lower-cost competitors
  • Current Arena rank is strong but outside the leading tier