Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Hermes 4.3 36B · Evaluations · Kaino
Hermes 4.3 36B logo
model evaluation

Hermes 4.3 36B

Nous Research

Open-weight 36B hybrid-reasoning language model from Nous Research with long-context support and local-deployment-oriented availability.

llmnous-research
65.7KAINO SCORENot recommended
Evaluated Jul 31, 20267 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness78
  • Developer experience78
  • Technical capability75
  • Reasoning & knowledge70
  • Risk & evidence68
  • Adoption signal62
  • Speed & availability58
  • Multimodal & I/O55
  • Coding & agentic55
  • Pricing clarity45

Kainotomic evaluation

Official Nous materials establish a 36B open-weight hybrid-reasoning model with 128K context, structured outputs, GGUF variants, and local-deployment support. Microsoft Foundry additionally documents a vLLM-backed, OpenAI-compatible Chat Completions deployment. These are practical strengths, but supplied evidence contains no independently reported quality, coding, latency, or preference result. Technical and reasoning scores therefore sit around the GLM-4.6 anchor rather than near DeepSeek-V3.2 or frontier hosted models. Coding is scored materially below DeepSeek-V3.2 (84) and Claude Opus 4.8 (95): DeepSWE does not list Hermes 4.3 36B, LiveCodeBench has no matching entry, and no SWE-bench result was found. Developer experience exceeds GLM-4.6's 73 on the documented GGUF and OpenAI-compatible deployment paths, but remains below Gemma 3's 86 because the supplied record provides less ecosystem and workflow evidence. Multimodal I/O is below Gemma 3 because official materials here substantiate text capabilities only. Open weights can make cost attractive for operators able to self-host, but 36B inference hardware requirements and absent hosted rates limit that advantage. Pricing clarity and availability are weak: Nous Portal marks the model unsupported, unavailable through its API, and priced N/A. Public adoption is modestly scored from its established distribution channels, not benchmark or usage statistics. Evidence quality is constrained by missing independent evaluations and an unspecified license.

Strengths

  • Open-weight 36B model with documented GGUF variants and local-deployment orientation
  • 128K context, hybrid-thinking positioning, and structured-output support in official materials
  • Microsoft Foundry documents vLLM-backed OpenAI-compatible Chat Completions

Caveats

  • No direct DeepSWE, SWE-bench, LiveCodeBench, Terminal-Bench, Aider, or preference result was supplied
  • Official evidence supports text generation and reasoning claims, not multimodal capability
  • Nous API availability is explicitly absent and no hosted price is supplied