Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
MiniMax M2.5 · Evaluations · Kaino
MiniMax M2.5 logo
model evaluation

MiniMax M2.5

MiniMax

Coding- and agent-focused MiniMax text model with reasoning, tool use, high-throughput API variants, and open weights for local deployment.

modelllmMiniMax
75.3KAINO SCORERecommended
Evaluated Jul 31, 202611 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness88
  • Coding & agentic84
  • Speed & availability82
  • Developer experience80
  • Technical capability79
  • Pricing clarity74
  • Risk & evidence72
  • Reasoning & knowledge71
  • Adoption signal68
  • Multimodal & I/O55

Kainotomic evaluation

MiniMax M2.5 has credible coding-oriented positioning, tool use, hosted API support, and open-weight deployment options in official MiniMax materials. Its reported 80.2% SWE-bench Verified result supports a coding score above GLM-4.6 (79), although it is vendor-reported, used internal infrastructure and Claude Code, and is not directly comparable to independently standardized runs. It therefore remains below DeepSeek-V3.2 (84), Kimi K2.5 (86), and frontier OpenAI/Anthropic anchors (95–96) for coding and agents. Artificial Analysis reports a 34 Intelligence Index, 75.3 output tokens/s, 1.70s TTFT, and $0.30/$1.20 per million input/output tokens. Those figures support strong value and above-median speed relative to the calibration set, though not DeepSeek-V3.2’s cost leadership. The supplied evidence establishes a text-focused model rather than broad native multimodal capability, so multimodal/I-O is held at the text-model floor. Open weights and API documentation improve developer experience, but supplied sources do not establish license terms or mature ecosystem depth. Independent coverage is incomplete: M2.5 is absent from the checked DeepSWE and LiveCodeBench leaderboards, and no Terminal-Bench or Aider result is supplied. On Arena Hard Prompts it ranks 143rd at 1415±5 from 25,350 votes, indicating meaningful exposure but modest preference relative to leading models. This evidence profile warrants conservative reasoning, adoption, and evidence-quality scores despite favorable vendor claims.

Strengths

  • Vendor-reported 80.2% SWE-bench Verified result indicates substantial repository-level coding capability.
  • Low reported API pricing and high measured output speed offer strong hosted value.
  • Official API support plus open-weight local deployment provide practical deployment flexibility.

Caveats

  • SWE-bench result is vendor-reported and run on internal infrastructure with Claude Code; it should not be treated as an independently reproduced ranking.
  • No supplied evidence demonstrates native image, audio, or video input/output capability.
  • Arena Hard preference rank is low relative to leading general-purpose chat models.