MODELS
by Xiaomi
Native omni-modal agent foundation model from Xiaomi MiMo with a 1M-token context window for image, video, audio, and text understanding.
MiMo-V2.5 is part of Xiaomi’s MiMo-V2.5 series. Xiaomi’s official MiMo homepage describes mimo-v2.5 as an omni-modal agent foundation model with a 1M context window and support for image, video, audio, and text understanding. Xiaomi’s MiMo API Open Platform states that the MiMo-V2.5 series, including mimo-v2.5 and mimo-v2.5-pro, was officially open-sourced under the MIT license. OpenRouter describes MiMo-V2.5 as a native omnimodal model by Xiaomi that delivers Pro-level agentic performance at roughly half the inference cost and surpasses MiMo-V2-Omni in multimodal perception across image and video understanding.
Use for omni-modal agent workflows that need image video audio and text understanding in one model. Use when a 1M-token context window is useful for long-context multimodal tasks. Evaluate when comparing Xiaomi MiMo-V2.5 against MiMo-V2-Omni or other multimodal foundation models.
Last updated May 31, 2026