Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
OpenHands LM · Evaluations · Kaino
O
model evaluation

OpenHands LM

All Hands AI / OpenHands

OpenHands LM 32B v0.1 is an open coding-agent model from OpenHands / All Hands AI for autonomous software-development tasks.

modelagentic-codingOpenHands
66.9KAINO SCORERecommended
Evaluated Jul 31, 20269 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Developer experience82
  • Cost effectiveness80
  • Technical capability74
  • Coding & agentic70
  • Risk & evidence68
  • Adoption signal68
  • Reasoning & knowledge65
  • Speed & availability62
  • Pricing clarity55
  • Multimodal & I/O45

Kainotomic evaluation

OpenHands LM is a focused, text-only 33B coding-agent model based on Qwen Coder 2.5 Instruct 32B. Its official launch material and model card report 37.2% SWE-bench Verified resolve rate and a 128K/131K context limit. That is meaningful task-specific evidence, but materially below the coding and broad-capability profiles represented by Claude Opus 4.8, GPT-5.5, GPT-5.6 Sol, and Kimi K2.5. It is also below DeepSeek-V3.2 and GLM-4.6 on the calibrated coding score because their published evaluations support stronger or broader current performance. MIT-licensed downloadable weights are its main practical advantage. Local deployment can make it cost-effective for teams with suitable hardware, and OpenHands provides a mature open workflow, documentation, runtime material, and repository integration. This supports a stronger developer-experience score than GLM-4.6 and DeepSeek-V3.2 in the supplied anchors. The supplied material does not establish image, audio, structured tool-I/O, throughput, latency, or a concrete hosted price, so multimodal, speed, and pricing scores remain conservative. Evidence quality is moderate rather than high: the SWE-bench figure is provider-reported, though repeated in the official model card. No DeepSWE/DataCurve result, LiveCodeBench score, Terminal-Bench/Aider result, or LMArena/Arena-Hard preference result is published in the supplied evidence. Artificial Analysis was checked but supplied without extractable benchmark, price, or speed figures. Adoption is supported by the OpenHands ecosystem, not model-specific usage data.

Strengths

  • Officially reported 37.2% SWE-bench Verified resolve rate for a dedicated open coding-agent model.
  • MIT-licensed downloadable BF16 weights enable local deployment.
  • 128K-class context and direct alignment with the documented OpenHands agent workflow.
  • Strong documentation and open-source ecosystem relative to similarly scored coding models.

Caveats

  • Text-generation-only evidence; no supported multimodal capability claim.
  • No supplied latency, throughput, or service-reliability measurements.
  • No exact model API price or total local-inference cost is established.
  • The reported SWE-bench result is first-party and not complemented by independent current benchmark results.