Quick Facts
- Microsoft released Phi-4-reasoning-vision-15B with 15 billion parameters, trained on 200 billion tokens versus competitors’ trillion-plus tokens
- The model scored 17% higher than Google’s gemma-3-12b-it on MathVista_Mini benchmark while using significantly less compute
- Training required only 240 NVIDIA B200 GPUs for 4 days, costing a fraction of typical multimodal model development
Microsoft released Phi-4-reasoning-vision-15B on March 4, 2026, an open-weight multimodal AI model that processes both text and images with 15 billion parameters. The model was trained on approximately 200 billion tokens of multimodal data, dramatically less than competitors.
Rival models from Alibaba’s Qwen family, Moonshot AI’s Kimi-VL, SenseTime’s InternVL series, and Google’s Gemma3 each consumed more than one trillion tokens during training. That’s roughly five times the total data pipeline Microsoft used.
Performance Metrics
On Microsoft’s internal evaluations across ten benchmarks, the model scored 84.8 on AI2D science diagrams, 83.3 on ChartQA, 75.2 on MathVista, 88.2 on ScreenSpot v2 for UI element grounding, and 54.3 on MMMU multimodal understanding tests.
The model scored 17% higher than Google’s gemma-3-12b-it on MathVista_Mini, a benchmark comprising multimodal math questions. These numbers generally trail much larger models like Qwen3-VL-32B but remain competitive with similarly-sized systems.
Technical Innovation
The model uses a mid-fusion architecture combining the Phi-4-Reasoning language model backbone with the SigLIP-2 vision encoder. It operates as a single system that can invoke extended chain-of-thought reasoning for mathematical and scientific tasks or default to direct inference for perception-focused tasks like captioning.
“We have competitive performance to much slower models that require ten times or more compute-time and tokens and better accuracy than similarly fast models, particularly when it comes to math and science reasoning,” Microsoft researchers wrote in a blog post.
Business Applications
The model targets AI agents that interact with applications via their user interfaces. With high-resolution perception capabilities, it can identify and localize interactive elements like buttons, menus, and text fields. This positioning aligns with industry predictions that autonomous software agents represent the next major AI frontier.
GlobalData anticipates 2026 will be the year of efficiency, with small language models gaining relevance for domain-specific enterprise use cases. Small models running on-device are cutting cloud costs by up to 70%.
Market Impact
Training large AI models costs millions in cloud compute, and the environmental footprint of trillion-token training runs draws increasing regulatory scrutiny. Microsoft’s approach delivers competitive results in a fraction of the time required by larger models.
GlobalData’s Agentic AI Forecast estimates global revenues will grow at a 48% CAGR from 2024 to 2029, with sales projected to reach $45.4 billion by 2029 from $6.4 billion in 2024.
The model is available under MIT license with a context length of 16,384 tokens and dynamic resolution vision encoder supporting up to 3,600 visual tokens.
Read more: Microsoft open-sources multimodal reasoning model with 15B parameters
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
