Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
GPT-5.6 Luna · Evaluations · Kaino
GPT-5.6 Luna logo
model evaluation

GPT-5.6 Luna

OpenAI

OpenAI’s fastest, lowest-cost GPT-5.6 model for cost-sensitive, high-volume text, coding, reasoning, and tool-use workloads via the OpenAI API.

modelopenaigpt-5.6-luna
78.8KAINO SCORERecommended
Evaluated Jul 30, 202610 reviews
Website Docs

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Pricing clarity90
  • Cost effectiveness86
  • Developer experience84
  • Speed & availability84
  • Multimodal & I/O82
  • Technical capability82
  • Coding & agentic78
  • Risk & evidence72
  • Reasoning & knowledge69
  • Adoption signal63

Kainotomic evaluation

OpenAI’s supplied API and product sources support GPT-5.6 Luna as a real API model positioned for cost-sensitive, high-volume use, with text and image input, text output, Responses API tools, a 1,050,000-token context window, and 128,000 max output tokens. Pricing is unusually clear in the supplied evidence: $1.00/M input, $0.10/M cached input, and $6.00/M output, with Artificial Analysis independently repeating the same price and 1.0M context figure. Coding evidence is positive but incomplete. DeepSWE/DataCurve lists gpt-5.6-luna[max] at 67%±4% over 113 tasks, with reported average cost, output-token, and step counts, suggesting capable agentic coding at moderate cost. LiveCodeBench direct search found no GPT-5.6/Luna listing, and Requesty marks LiveCodeBench as N/A while giving a separate Coding Index. No SWE-bench result was found. Terminal-Bench/Aider were checked, but the supplied sources do not provide a model-specific score. Public preference and general intelligence signals are mixed. Artificial Analysis reports Intelligence Index 46, high throughput at 202.6 output tokens/s, but TTFT latency of 7.83s. LMArena dataset evidence lists GPT 5.6 Luna xHigh at rank 18, while Prolific’s HUMAINE report places Luna #36 of 54. Overall, Luna looks like a strong low-cost OpenAI production option, but not a top frontier model; confidence is reduced by missing SWE-bench, LiveCodeBench, and Terminal-Bench/Aider results.

Strengths

  • Official OpenAI API evidence for model availability, context window, output limit, tools, and multimodal input
  • Clear token pricing from OpenAI with third-party corroboration from Artificial Analysis
  • Strong cost/speed profile for high-volume workloads
  • DeepSWE evidence indicates credible agentic coding performance

Caveats

  • LiveCodeBench leaderboard search found no direct GPT-5.6 Luna listing
  • No SWE-bench or SWE-bench Verified result was found
  • Terminal-Bench/Aider family was checked, but supplied evidence does not include a model-specific score
  • Public preference evidence is mixed rather than consistently high