Microsoft
Compact 15B open-weight multimodal reasoning model for vision-language tasks, UI grounding, math, science, and document understanding.
Phi-4-reasoning-vision-15B is a compact open-weight multimodal model with official positioning for vision-language reasoning, UI grounding, documents, math, and science. Its 15B scale, direct-versus-longer reasoning design, and documented vision focus support technical and multimodal scores above GLM-4.6’s 55 multimodal anchor, but below Kimi K2.5 (84) and Gemini 3.5 Flash (92), which have stronger evaluated multimodal ecosystems. The supplied evidence supports capability directionally, not a broad independent performance ranking. Coding and agentic-work evidence is weak: DeepSWE and LiveCodeBench explicitly have no listing, while no usable Terminal-Bench, Aider, or SWE-bench result is supplied. Accordingly it sits materially below GLM-4.6 (79), Grok 3 (82), and frontier coding anchors. Open weights and official Hugging Face/GitHub distribution improve deployability and likely self-hosting economics relative to closed premium models such as Claude Opus 4.8, but Microsoft pricing or serving-cost figures are not supplied. Developer experience is solid for an officially published repository and model page, though API, production-hosting, throughput, adoption, and safety evidence is limited. Speed receives only a moderate compact-model inference rather than a measured claim. Public signal remains below widely deployed proprietary anchors, and evidence quality is constrained by reliance on official materials plus benchmark absences.