
MODELS
Phi-4-reasoning-vision-15B
by Microsoft
Overview
Compact 15B open-weight multimodal reasoning model for vision-language tasks, UI grounding, math, science, and document understanding.
Details
Phi-4-reasoning-vision-15B is a Microsoft model described in the supplied sources as a compact 15B open-weight multimodal reasoning model. The referenced Microsoft Research page focuses on “Phi-4-reasoning-vision” and training lessons for a multimodal reasoning model, while the Hugging Face and GitHub URLs identify the specific Phi-4-reasoning-vision-15B model. The provided source excerpts position it for vision-language tasks, UI grounding, math, science, and document understanding.
When to Use
Use when evaluating a compact open-weight multimodal reasoning model from Microsoft for vision-language tasks. Use for experiments involving UI grounding, document understanding, math, or science tasks where the supplied model documentation and repository should be reviewed first.
Getting Started
- Read the Microsoft Research blog post for the model overview and training context.
- Open the Hugging Face model page for model-specific documentation and usage information.
- Review the Microsoft GitHub repository before running trials or integrating the model into an application.
Key Features
- •15B-parameter compact multimodal reasoning model, according to the supplied source excerpt.
- •Open-weight availability is stated in the supplied source excerpt.
- •Positioned for vision-language tasks, UI grounding, math, science, and document understanding.
Capabilities
- •multimodal reasoning
- •vision-language tasks
- •UI grounding
- •document understanding
Last updated Jul 31, 2026