
MODELS
by StepFun
High-efficiency multimodal sparse MoE vision-language model from StepFun, available through StepFun API and chat interfaces.
Step 3.7 Flash is StepFun’s high-efficiency multimodal Mixture-of-Experts model. Source descriptions identify it as a sparse MoE vision-language model with a roughly 196B/198B-parameter scale, native image and video understanding through a vision encoder, around 11B active parameters, a 256k context window, selectable reasoning levels, and local deployment instructions in the official GitHub repository. StepFun’s official page says it is available through StepFun’s API platform and web/app chat interfaces.
Use when you need a StepFun multimodal model for text, image, or video understanding workflows. Use when long-context tasks may benefit from the model’s stated 256k context window. Use when you want selectable reasoning levels or a high-efficiency sparse MoE vision-language model to evaluate.
Last updated May 29, 2026