Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Muse Spark · Evaluations · Kaino
Muse Spark logo
model evaluation

Muse Spark

Meta

Muse Spark is a Meta Superintelligence Labs model family described by Meta as focused on multimodal reasoning, tool use, coding, and agentic tasks.

metamuse-spark
82.1KAINO SCORERecommended
Evaluated Jul 31, 202613 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Multimodal & I/O87
  • Cost effectiveness87
  • Developer experience84
  • Technical capability84
  • Speed & availability84
  • Coding & agentic84
  • Pricing clarity82
  • Reasoning & knowledge79
  • Risk & evidence78
  • Adoption signal72

Kainotomic evaluation

Muse Spark 1.1 has credible high-end capability evidence, but not evidence of leadership. Meta reports 77.4% SWE-Bench Verified under a 15-attempt methodology; LiveCodeBench tracking reports 69.7% (15th/72), while DeepSWE reports 53% ±3% (11th/18). This supports coding/agentic performance around DeepSeek-V3.2 and below Kimi K2.5, GPT-5.5, GPT-5.6 Sol, and Claude Opus 4.8. Meta’s own evaluation says it trails Opus 4.8 and/or GPT-5.5 on Terminal-Bench 2.1 and SWE-Bench Pro. Multimodal and API breadth are relative strengths: official material documents native multimodal reasoning, tool/function calling, and a managed 1M-token context. Its 51 Artificial Analysis Intelligence Index and 1505±7 Arena hard-prompts result (14th, preliminary; 8,670 votes) indicate solid general reasoning and preference, not frontier placement. The $1.25/M input and $4.25/M output pricing is materially more economical than premium frontier anchors, while AA’s 127.3 tokens/s and 2.47s TTFT support a top-tier speed score. Public preview status tempers availability and adoption. Developer experience is above GLM-4.6 and DeepSeek-V3.2 due to Meta-hosted API access and agentic affordances, but below GPT-5.5’s more mature tooling ecosystem. Risk/evidence quality is relatively strong: Meta provides evaluation, safety, and preparedness reporting, yet key performance claims remain substantially vendor-reported and several independent benchmark families provide limited or no directly comparable detail. Pricing is independently listed, but official pricing and commercial terms were not supplied.

Strengths

  • Documented multimodal reasoning, tool/function calling, and 1M-token context
  • Competitive coding evidence across SWE-Bench Verified, LiveCodeBench, and DeepSWE
  • Low listed token pricing with strong reported throughput
  • Published evaluation and safety/preparedness materials

Caveats

  • DeepSWE rank is 11th of 18 and Meta reports trailing leading proprietary models on Terminal-Bench 2.1 and/or SWE-Bench Pro
  • Arena result is preliminary preference evidence rather than a capability benchmark
  • Public-preview positioning and limited ecosystem evidence constrain adoption and availability assessments
  • Official pricing and commercial terms were not provided in the supplied documentation