Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Morph V3 Large · Evaluations · Kaino
M
model evaluation

Morph V3 Large

Morph

Specialized code-apply model for merging AI-generated code edits with high accuracy, aimed at coding agents and developer-product workflows.

modelsource:morphllm.comcodecode-applycoding-agentsdeveloper-tools
66.6KAINO SCORENot recommended
Evaluated Jul 31, 20266 reviews
Website Docs

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Pricing clarity89
  • Cost effectiveness89
  • Speed & availability82
  • Developer experience78
  • Technical capability72
  • Coding & agentic65
  • Risk & evidence61
  • Adoption signal55
  • Reasoning & knowledge45
  • Multimodal & I/O30

Kainotomic evaluation

Morph V3 Large is a narrowly specialized apply/merge model, not a demonstrated general-purpose coding or reasoning model. Official documentation supports its intended code-edit application workflow, and Morph lists a 262K context window. Its technical and coding-agentic scores sit below broad, benchmarked systems such as Claude Opus 4.8, GPT-5.5, GPT-5.6 Sol, and Kimi K2.5: those anchors have materially stronger evidence for end-to-end software work, whereas Morph has no reported DeepSWE, SWE-bench, or LiveCodeBench result. The commercial case is unusually favorable on the supplied provider pricing: $0.90/M input and $1.90/M output tokens, alongside a claimed 5,000+ tokens/sec. This supports higher cost-effectiveness and speed scores than costly frontier anchors such as Opus 4.8 and GPT-5.5, while the explicit per-token rates also make pricing clearer than Hermes 4.3 36B’s published anchor. Those performance figures are provider claims rather than independently verified measurements. Developer experience is credible for the discrete apply step because official product and model documentation exist, but capability boundaries matter: there is no supplied evidence for image/audio I/O, broad knowledge, autonomous repository resolution, public preference, or ecosystem adoption. The score is therefore closer to a focused infrastructure component than to GPT-5.6 Luna or Gemini 3.5 Flash as a full developer model. Public benchmark absence is a material evidence limitation, not evidence of poor code-application accuracy.

Strengths

  • Purpose-built positioning for applying and merging AI-generated code edits.
  • Published token pricing is low and explicit: $0.90/M input and $1.90/M output.
  • Official pricing claims 5,000+ tokens/sec and a 262K-token context window.
  • Official product and model documentation are available.

Caveats

  • Not evidenced as a general coding, planning, reasoning, or autonomous-agent model.
  • No supplied evidence of multimodal inputs or outputs.
  • Speed and accuracy claims are provider-reported rather than independently validated.
  • The pricing source does not establish independently measured availability, latency, or reliability.