Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Granite 4.0 H Small · Evaluations · Kaino
Granite 4.0 H Small logo
model evaluation

Granite 4.0 H Small

IBM

IBM Granite-4.0-H-Small is a 32B total, 9B active hybrid MoE instruction model for enterprise RAG, agents, coding, tool use, and multilingual dialog.

modelibmgranite-4
73.3KAINO SCORERecommended
Evaluated Jul 31, 202612 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness89
  • Pricing clarity87
  • Developer experience83
  • Technical capability76
  • Speed & availability74
  • Risk & evidence73
  • Adoption signal68
  • Reasoning & knowledge67
  • Multimodal & I/O58
  • Coding & agentic58

Kainotomic evaluation

Granite 4.0 H Small has credible enterprise-oriented foundations: a 32B-total/9B-active hybrid MoE design, 131k context, instruction following, RAG, function calling, tool use, FIM coding, and multilingual dialog are documented by IBM. Its Apache-2.0 release and Hugging Face availability make it materially more deployable than closed API-only options. However, Artificial Analysis reports an Intelligence Index of 5, and the public Arena Hard result—rank 256, 1240±11 across 2,993 votes—does not support placing general capability or reasoning near Gemini 3.5 Flash, Kimi K2.5, GPT-5.5, or Claude Opus 4.8. Coding and agent-work scores remain conservative despite the documented tools and code features. DeepSWE and LiveCodeBench explicitly do not list the model; the supplied SWE-bench and Terminal-Bench/Aider materials provide no model-specific published result. This leaves it below DeepSeek-V3.2 and Kimi K2.5 on demonstrated agentic coding, while above Gemma 3’s low coding anchor only modestly on product support rather than benchmark proof. The model is text-centric in the supplied evidence, so its I/O score is closer to DeepSeek-V3.2 than multimodal Gemini or Claude offerings. Cost is the clearest advantage: IBM lists $0.0636/M input and $0.265/M output tokens, consistent with Artificial Analysis estimates, substantially under premium anchors. Artificial Analysis records high throughput (386.4 tok/s), but 10.22s TTFT and only one tracked provider limit the availability score. IBM’s documentation, watsonx API integration, open weights, and clear pricing support a strong developer score; limited independent quality evidence, a recent release, and weak preference standing constrain adoption and evidence-confidence scores.

Strengths

  • Documented 131k context, tool/function calling, RAG, FIM code, and multilingual dialog
  • Apache-2.0 model release with Hugging Face and watsonx.ai access
  • Very low published token pricing and high measured output throughput

Caveats

  • Artificial Analysis Intelligence Index of 5 and Arena Hard rank 256 indicate limited demonstrated general capability
  • Text-focused supplied evidence does not establish vision, audio, or broad multimodal performance
  • 10.22s measured TTFT and one tracked Artificial Analysis provider temper serving claims