Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Qwen3-Coder-Next · Evaluations · Kaino
Qwen3-Coder-Next logo
model evaluation

Qwen3-Coder-Next

Alibaba Qwen

Open-weight Qwen coding-agent model for software engineering, tool calling, structured outputs, and long-context developer workflows.

modelalibaba-qwen
79.2KAINO SCORERecommended
Evaluated Jul 31, 202612 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness88
  • Developer experience87
  • Technical capability84
  • Speed & availability84
  • Coding & agentic84
  • Pricing clarity78
  • Risk & evidence77
  • Reasoning & knowledge75
  • Adoption signal75
  • Multimodal & I/O60

Kainotomic evaluation

Qwen3-Coder-Next has strong, model-specific coding evidence: its technical report reports 70.6%–71.3% on SWE-bench Verified depending on the agent scaffold, while CodeSOTA lists 58.93 pass@1 and rank 20 on LiveCodeBench. This supports a coding/agentic score above Grok 3’s published 82, but below Kimi K2.5 (86) and frontier OpenAI/Anthropic coding anchors (95–96), which have stronger comparative evidence. The 80B-total/3B-active MoE design, 262K context, tool calling, and executable-environment focus are meaningful technical strengths. General reasoning evidence is thinner than coding evidence, and no scored multimodal result was supplied. Cost and serving are unusually favorable relative to the published anchors. Artificial Analysis reports $0.35/M input and $1.20/M output median pricing, 106.8 tokens/s median throughput, and 1.70s TTFT; tracked providers show faster options. That justifies cost effectiveness above Gemini 3.5 Flash (76) and speed near the catalog high range. Apache-2.0 weights, deployment guidance, OpenAI-compatible Model Studio access, and a substantial Qwen GitHub ecosystem support a developer-experience score above most closed-model anchors. Pricing clarity is lower than the most transparent API anchors because Model Studio tiers, third-party routes, and self-hosting produce materially different costs. Evidence quality is good but not frontier-grade: official material and a technical report substantiate core claims, while several benchmark references lack independently reported scores. DeepSWE has no matching entry. TerminalBench/Aider and Arena-Hard were checked, but supplied sources only establish that they were considered, not a result. Public adoption is credible through Qwen’s distribution and multi-provider availability, but no direct usage, download, or preference ranking was supplied.

Strengths

  • SWE-bench Verified results of 70.6%–71.3% across reported agent scaffolds
  • 80B-total/3B-active Apache-2.0 MoE with 262K context
  • Low tracked API pricing and high measured provider throughput
  • Open-weight deployment plus OpenAI-compatible hosted access

Caveats

  • LiveCodeBench rank 20/pass@1 58.93 indicates solid rather than leading standalone code generation
  • No supplied scored evidence for multimodal capability, TerminalBench, Aider, Arena-Hard, or LMArena
  • Benchmarked SWE performance depends on the selected agent scaffold
  • Prices and latency vary substantially across Model Studio and third-party providers