Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Sarvam-105B · Evaluations · Kaino
S
model evaluation

Sarvam-105B

Sarvam AI

Open-source 105B-parameter MoE reasoning chat model with an OpenAI-compatible API, long-context support, coding, agentic, and Indian-language strengths.

open-sourcecodingopenai-compatible-apisarvam-ai
67.4KAINO SCORENot recommended
Evaluated Jul 31, 202612 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness88
  • Developer experience76
  • Technical capability75
  • Reasoning & knowledge72
  • Risk & evidence70
  • Pricing clarity68
  • Adoption signal65
  • Multimodal & I/O55
  • Coding & agentic53
  • Speed & availability52

Kainotomic evaluation

Sarvam-105B is a 105B MoE open-weights reasoning model with a documented OpenAI-compatible API, 128k context, and stated Indian-language specialization. Its $0.042/M input and $0.17/M output estimate is materially more economical than most frontier hosted models and supports a cost score above Grok 3 (70) and near DeepSeek-V3.2 (91). However, Artificial Analysis’ Intelligence Index of 12 and the absence of independent broad capability results keep technical capability below DeepSeek-V3.2 (82), Kimi K2.5 (83), and frontier anchors. Coding evidence is mixed. The official model card reports 45.0 on SWE-Bench Verified using SWE-Agent, but Artificial Analysis reports only 1.5% on TerminalBench Hard; this supports a coding/agentic score below Hermes 4.3 36B (55) and far below GLM-4.6 (79) or DeepSeek-V3.2 (84). A 0.710 Arena-Hard-v2 score at rank 10 provides some preference-based support for general reasoning, but it is not a substitute for independently replicated knowledge or reasoning benchmarks. No confirmed LiveCodeBench leaderboard entry was retrieved. Developer usability benefits from open weights, Hugging Face distribution, and OpenAI-compatible access, placing it slightly above GLM-4.6 on interface maturity. Multimodal capability is not evidenced in supplied materials. Only Sarvam is tracked as an API provider, and no measured throughput, latency, or response-time data is available, limiting availability and speed confidence. Official benchmark reporting is useful but should not outweigh the sparse independent evidence.

Strengths

  • Open-weights 105B MoE model with OpenAI-compatible API and 128k context.
  • Very low tracked token pricing relative to hosted frontier models.
  • Documented focus on Indian-language use cases.
  • Official SWE-Bench Verified result and public Arena-Hard-v2 listing provide limited performance evidence.

Caveats

  • TerminalBench Hard result of 1.5% is weak evidence for autonomous terminal-agent work.
  • No independently retrieved LiveCodeBench entry confirms the provider-reported figure.
  • No supplied evidence supports image, audio, video, or other multimodal I/O.
  • Speed and latency are unmeasured, with one tracked API provider.