Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Step 3.5 Flash · Evaluations · Kaino
S
model evaluation

Step 3.5 Flash

StepFun

Open sparse MoE reasoning model from StepFun with 196B total parameters, 11B active parameters, 256K context, and support for coding, tool use, agentic workflows, deep research, API access, and local deployment.

agenticstepfun
77.2KAINO SCORERecommended
Evaluated Jul 31, 202613 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness88
  • Pricing clarity86
  • Speed & availability84
  • Coding & agentic82
  • Developer experience81
  • Technical capability80
  • Reasoning & knowledge74
  • Risk & evidence73
  • Adoption signal69
  • Multimodal & I/O55

Kainotomic evaluation

Step 3.5 Flash has credible capability evidence for an efficient open MoE: 196B total/11B active parameters, 256K context, API access, and local weights. Artificial Analysis assigns an Intelligence Index of 26, placing it below frontier anchors such as GPT-5.5 and Claude Opus 4.8, but its measured 303.3 tokens/s and 1.05s TTFT support a higher speed score than DeepSeek-V3.2. Multimodal support is not substantiated in the supplied evidence, so this remains comparable to the low-I/O GLM-4.6 and DeepSeek anchor scores rather than Kimi K2.5. Coding evidence is promising but partly provider-reported: StepFun’s model card reports 74.4% SWE-bench Verified and Terminal-Bench 2 reports 51.0; the indexed 86.4% LiveCodeBench V6 result is explicitly self-reported. This supports coding/agentic performance above GLM-4.6 and far above Gemma 3, while remaining below Kimi K2.5 and frontier OpenAI/Anthropic anchors with stronger independent breadth. Reasoning is moderated by a 74.0 vendor-published Arena-Hard-v2 result and a 1414±4 Text Arena Hard score/rank 146 from 35,083 votes. At $0.10/M input and $0.30/M output, it is markedly more economical than premium frontier anchors and approaches DeepSeek-V3.2’s cost position. Official API, pricing, GitHub, and Hugging Face availability improve developer usability, though the original AA listing is deprecated and license terms were not established. Public adoption is material but below established leading providers. Evidence quality is reduced by vendor-origin benchmark reporting and no DeepSWE leaderboard entry.

Strengths

  • Very high measured generation speed and low latency in Artificial Analysis.
  • Low documented API pricing with cached-input pricing and tiered limits.
  • Open-weight distribution, API access, 256K context, and documented tool-use/coding positioning.
  • Reported SWE-bench Verified and Terminal-Bench 2 results provide concrete coding evidence.

Caveats

  • LiveCodeBench V6 result is indexed as self-reported.
  • No supplied evidence establishes multimodal input/output support.
  • Original Step 3.5 Flash is labeled deprecated by Artificial Analysis; the 2603 variant may differ.
  • Terminal-Bench and SWE-bench results originate from StepFun-controlled sources rather than an independent leaderboard.