Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Mistral Medium 3.5 · Evaluations · Kaino
Mistral Medium 3.5 logo
model evaluation

Mistral Medium 3.5

Mistral AI

Frontier-class Mistral model optimized for agentic and coding use cases, with reasoning support and multimodal capabilities for developers.

modelmistral-ai
77.1KAINO SCORERecommended
Evaluated Jul 31, 202612 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Developer experience84
  • Speed & availability83
  • Coding & agentic80
  • Technical capability79
  • Multimodal & I/O78
  • Cost effectiveness78
  • Pricing clarity74
  • Risk & evidence73
  • Reasoning & knowledge72
  • Adoption signal70

Kainotomic evaluation

Mistral Medium 3.5 has credible upper-mid-tier developer capability: Mistral’s card identifies a frontier-class multimodal model for coding and agents, while Artificial Analysis reports a 30 Intelligence Index and a 256k context window. Its 77.6% SWE-bench Verified result is vendor-reported, not an independently reproduced leaderboard result. This places it below GPT-5.5 and Claude Opus 4.8 on demonstrated frontier capability, but above GLM-4.6’s technical baseline and close to the lower end of Kimi K2.5’s range. Coding evidence is mixed but usable. The supplied Terminal-Bench aggregator reports 50.6% on Terminal-Bench 2.1 (rank 54/142), supporting competent terminal work but not the leading agentic tier; DeepSWE and LiveCodeBench provide no listed result. Accordingly it trails DeepSeek-V3.2 and Kimi K2.5 in coding/agentic confidence despite the SWE-bench claim. Official multimodal positioning supports a materially higher I/O score than text-oriented DeepSeek-V3.2 and GLM-4.6, though modality-level quality details are absent. At $1.50/M input and $7.50/M output, it is materially cheaper than premium frontier anchors but not a budget leader. Artificial Analysis’ 76.3 output tok/s and 2.16s TTFT support an above-median speed score. Mistral documentation and API materials support solid developer experience. Public preference is moderate: Text Arena rank 87, 1427±7, across 11,019 votes shows meaningful usage but substantially weaker preference than leading published anchors.

Strengths

  • Officially documented multimodal, coding, reasoning, and agentic positioning
  • Vendor-reported 77.6% SWE-bench Verified result
  • 256k context window and strong measured throughput
  • Mid-market token pricing relative to premium frontier models

Caveats

  • SWE-bench figure is provider-reported in the supplied evidence
  • Terminal-Bench evidence comes from BenchmarkList rather than the primary benchmark
  • No listed DeepSWE or LiveCodeBench result
  • Arena preference rank is modest despite substantial vote volume