Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
gpt-oss-20b · Evaluations · Kaino
gpt-oss-20b logo
model evaluation

gpt-oss-20b

OpenAI

OpenAI open-weight 20B model for local and developer workflows, including coding, reasoning, and tool-use experiments.

model-20bOpenAI
75.4KAINO SCORERecommended
Evaluated Jul 31, 202613 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness91
  • Developer experience87
  • Speed & availability84
  • Pricing clarity83
  • Technical capability76
  • Risk & evidence76
  • Coding & agentic68
  • Reasoning & knowledge67
  • Adoption signal67
  • Multimodal & I/O55

Kainotomic evaluation

gpt-oss-20b is a compact open-weight reasoning model with official documentation and repository support for local deployment, coding, tool use, and developer experimentation. Its 131k context window, reported 200.6 output tokens/s, and $0.06/$0.20 per-million-token hosted pricing in Artificial Analysis support unusually strong cost and speed scores. The supplied evidence does not establish vision or broad multimodal input, so multimodal capability remains at the calibration floor. Coding evidence is mixed. OpenAI reports SWE-bench Verified results rising from 37.4% at low to 60.7% at high reasoning, plus 2,230 Codeforces Elo without tools. However, Terminal-Bench 2.0 reports 3% ±1 across 74 tasks with Mini-SWE-Agent and Terminus 2, and DeepSWE and LiveCodeBench publish no result. This supports capability above Gemma 3's coding anchor (42), but below DeepSeek-V3.2 (84) and Kimi K2.5 (86), whose catalog scores reflect stronger agentic evidence. Reasoning and public-preference evidence is modest: Arena lists rank 231 overall and the cited third-party Arena-Hard V2 page reports 48.6%, well below frontier anchors such as Claude Opus 4.8 and GPT-5.5. Open weights, official setup materials, and low serving cost make developer experience strong. Pricing is relatively clear through the cited external rate card, but provider-controlled terms, deployment availability, and license details should be verified for a chosen host.

Strengths

  • Open-weight local deployment with official OpenAI documentation and repository
  • High-reasoning SWE-bench Verified result of 60.7%
  • Low reported hosted token pricing and high reported output throughput
  • Long 131k context window and documented developer workflow positioning

Caveats

  • Terminal-Bench 2.0 result is 3% ±1, indicating weak demonstrated autonomous terminal-agent performance
  • No published gpt-oss-20b entry was found on DeepSWE or LiveCodeBench
  • Arena rank 231 and reported 48.6% Arena-Hard V2 indicate limited general-preference competitiveness
  • No supplied evidence establishes image or other multimodal input support