Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
GPT-5 · Evaluations · Kaino
GPT-5 logo
model evaluation

GPT-5

OpenAI

OpenAI model for coding and agentic tasks with text and image input, structured outputs, and parallel tool calling.

OpenAIagentsGPT-5
81.4KAINO SCORERecommended
Evaluated Sep 7, 202611 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Developer experience91
  • Coding & agentic91
  • Technical capability89
  • Multimodal & I/O84
  • Reasoning & knowledge84
  • Adoption signal82
  • Pricing clarity78
  • Risk & evidence78
  • Cost effectiveness70
  • Speed & availability67

Kainotomic evaluation

GPT-5 has strong documented platform capability: a 400k-token context window, image and text input, structured outputs, function calling, parallel tool calls, streaming, and up to 128k output tokens. Its 74.9% SWE-bench Verified launch result and 74% system-card result on a fixed 477-task subset are substantial, though provider-reported. Independent leaderboard evidence is also strong: LiveCodeBench V6 reports 89.6%, while Aider Polyglot reports 88.0% with diff edits. Relative to GPT-5.5 and GPT-5.6 Sol, GPT-5 scores lower on technical capability, coding, and reasoning because the later models have stronger calibrated evidence, including DeepSWE entries for 5.5 and Sol. It scores materially above GPT-5.6 Terra in coding and developer experience because the exact GPT-5 has direct SWE-bench, LiveCodeBench, and Aider evidence plus mature API tooling. Its multimodal and I/O score is comparable to Kimi K2.5, but below Claude Opus 4.8 due to narrower supplied evidence on image-task quality. At $1.25/M input and $10/M output tokens in Artificial Analysis, cost is moderate rather than leading; the reported 69.93-second time to first token also constrains interactive speed despite 105.2 output tokens/s. Official documentation and a system card improve evidence quality, but several supplied records are future-dated and pricing is not directly substantiated by the catalog’s official-source extract. Arena confirms GPT-5’s presence, while the current ranking cited is for a later GPT-5-family variant, not this exact model.

Strengths

  • Direct exact-model coding evidence across SWE-bench Verified, LiveCodeBench V6, and Aider Polyglot.
  • Large context/output limits and mature agent APIs, including parallel tool calling and structured outputs.
  • Broad developer availability and strong public recognition for the GPT-5 family.

Caveats

  • SWE-bench figures are provider-reported and use different evaluation setups.
  • DeepSWE has no exact GPT-5 entry; results for GPT-5.5 and GPT-5.6 Sol cannot be attributed to GPT-5.
  • Reported high time to first token weakens latency-sensitive use cases.