Phi-4-reasoning-vision-15B logo

MODELS

Phi-4-reasoning-vision-15B

by Microsoft

modelsource:microsoft.commultimodalvision-languagereasoningopen-weightMicrosoft

Overview

Compact 15B open-weight multimodal reasoning model for vision-language tasks, UI grounding, math, science, and document understanding.

Details

Phi-4-reasoning-vision-15B is a Microsoft model described in the supplied sources as a compact 15B open-weight multimodal reasoning model. The referenced Microsoft Research page focuses on “Phi-4-reasoning-vision” and training lessons for a multimodal reasoning model, while the Hugging Face and GitHub URLs identify the specific Phi-4-reasoning-vision-15B model. The provided source excerpts position it for vision-language tasks, UI grounding, math, science, and document understanding.

When to Use

Use when evaluating a compact open-weight multimodal reasoning model from Microsoft for vision-language tasks. Use for experiments involving UI grounding, document understanding, math, or science tasks where the supplied model documentation and repository should be reviewed first.

Getting Started

  1. Read the Microsoft Research blog post for the model overview and training context.
  2. Open the Hugging Face model page for model-specific documentation and usage information.
  3. Review the Microsoft GitHub repository before running trials or integrating the model into an application.

Key Features

  • 15B-parameter compact multimodal reasoning model, according to the supplied source excerpt.
  • Open-weight availability is stated in the supplied source excerpt.
  • Positioned for vision-language tasks, UI grounding, math, science, and document understanding.

Capabilities

  • multimodal reasoning
  • vision-language tasks
  • UI grounding
  • document understanding

Last updated Jul 31, 2026