GPT-5 has strong documented platform capability: a 400k-token context window, image and text input, structured outputs, function calling, parallel tool calls, streaming, and up to 128k output tokens. Its 74.9% SWE-bench Verified launch result and 74% system-card result on a fixed 477-task subset are substantial, though provider-reported. Independent leaderboard evidence is also strong: LiveCodeBench V6 reports 89.6%, while Aider Polyglot reports 88.0% with diff edits. Relative to GPT-5.5 and GPT-5.6 Sol, GPT-5 scores lower on technical capability, coding, and reasoning because the later models have stronger calibrated evidence, including DeepSWE entries for 5.5 and Sol. It scores materially above GPT-5.6 Terra in coding and developer experience because the exact GPT-5 has direct SWE-bench, LiveCodeBench, and Aider evidence plus mature API tooling. Its multimodal and I/O score is comparable to Kimi K2.5, but below Claude Opus 4.8 due to narrower supplied evidence on image-task quality. At $1.25/M input and $10/M output tokens in Artificial Analysis, cost is moderate rather than leading; the reported 69.93-second time to first token also constrains interactive speed despite 105.2 output tokens/s. Official documentation and a system card improve evidence quality, but several supplied records are future-dated and pricing is not directly substantiated by the catalog’s official-source extract. Arena confirms GPT-5’s presence, while the current ranking cited is for a later GPT-5-family variant, not this exact model.