Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Llama 4 Maverick · Evaluations · Kaino
Llama 4 Maverick logo
model evaluation

Llama 4 Maverick

Meta

Meta Llama 4 Maverick is an open-weight, natively multimodal Mixture-of-Experts model for text and image understanding.

multimodalllamameta
79.6KAINO SCORERecommended
Evaluated Jul 31, 202614 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness91
  • Pricing clarity85
  • Multimodal & I/O84
  • Developer experience84
  • Speed & availability84
  • Technical capability81
  • Adoption signal78
  • Risk & evidence77
  • Reasoning & knowledge74
  • Coding & agentic58

Kainotomic evaluation

Maverick is a capable open-weight multimodal MoE: 17B active/400B total parameters, image-text understanding, and a documented 1M-token context. Its technical and multimodal scores sit above Gemma 3’s overall capability profile but below Gemini 3.5 Flash, whose evaluated I/O capability is stronger. Artificial Analysis reports very low median API prices and 105.3 output tokens/s, supporting unusually high cost-effectiveness and speed scores; Google Cloud’s published token rates also make pricing comparatively clear. Coding evidence is materially weaker than leading agent models. SWE-rebench reports 16.0% resolved on SWE-bench Verified, Terminal-Bench lists 15.5% ±3.3, and Meta reports 43.4% zero-shot LiveCodeBench pass@1 for a specified window. This supports a score above Gemma 3’s low coding anchor, but well below Kimi K2.5 and Claude Opus 4.8. The 1417 LMArena result is meaningful preference evidence, but it applies to an experimental chat variant rather than conclusively to the released instruct model. Meta supplies model cards, prompt guidance, safeguards, acceptable-use policy, source repository, and managed Bedrock access, giving solid developer support. Evidence quality is tempered by vendor-reported LiveCodeBench and preference claims, benchmark variant differences, and no listed DeepSWE or current LiveCodeBench leaderboard entry. Terminal-Bench and SWE-rebench provide useful independent checks, but they do not indicate frontier autonomous software-engineering performance.

Strengths

  • Open-weight natively multimodal MoE with documented 1M-token context
  • Very low published API pricing and strong independently reported throughput
  • Official model card, repository, prompt documentation, safety material, and managed-cloud access

Caveats

  • SWE-bench Verified and Terminal-Bench results are far below leading coding-agent anchors
  • LMArena 1417 applies to an experimental chat variant, not necessarily the released instruct model
  • Managed pricing and availability vary by provider despite clear Google Cloud rates