Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
EVI 4-mini · Evaluations · Kaino
E
model evaluation

EVI 4-mini

Hume AI

Multilingual real-time speech-language model for emotionally intelligent voice interfaces with streaming speech and prosody-aware responses.

modelvoice-aievi
64.6KAINO SCORENot recommended
Evaluated Jul 31, 20267 reviews
Website Docs

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Multimodal & I/O86
  • Pricing clarity84
  • Developer experience82
  • Cost effectiveness78
  • Technical capability74
  • Speed & availability72
  • Risk & evidence68
  • Adoption signal54
  • Reasoning & knowledge30
  • Coding & agentic18

Kainotomic evaluation

EVI 4-mini is a specialized real-time speech-to-speech interface model rather than a general foundation model. Official Hume documentation supports multilingual streaming voice interaction, prosody-aware behavior, and lower-latency positioning. Its multimodal/I-O score is consequently above text-first Grok 3 (76) and comparable to stronger voice-capable integrations, but below Gemini 3.5 Flash (92), for which the calibration record indicates broader demonstrated multimodal capability. Hume’s documentation and API focus support a solid developer-experience score. The model scores far below Gemma 3 (42), Kimi K2.5 (86), and Claude Opus 4.8 (95) in coding and agentic work because EVI 4-mini is not a coding model and has no DeepSWE, SWE-bench, LiveCodeBench, Terminal-Bench, or Aider result. Reasoning and knowledge are also limited: Hume explicitly states that EVI 4-mini does not natively generate language and requires a supplemental LLM. End-to-end assistant quality therefore depends materially on the paired model, orchestration, prompts, and speech pipeline. Published per-minute EVI tier rates ($0.04-$0.07) and included usage make pricing comparatively clear and potentially cost-effective for voice workloads, though no independent price-performance comparison was supplied. Lower latency is provider-stated, not independently benchmarked, so speed remains moderate. Public adoption and independent evidence are limited: no relevant leaderboard placement or preference evidence was found. Hume also cautions that emotion outputs are not direct inferences of a person’s feelings; applications should avoid consequential affect judgments and follow its stated ethical-use guidance.

Strengths

  • Native real-time, streaming speech-to-speech API design
  • Multilingual and prosody-aware interaction capabilities documented by Hume
  • Clear published per-minute usage rates and tier structure
  • Focused API documentation for voice-interface implementation

Caveats

  • Requires an external LLM for language generation
  • No direct coding, reasoning, or agent benchmark result was supplied
  • Latency and naturalness claims are provider assertions rather than independent measurements
  • End-to-end quality cannot be attributed to EVI 4-mini independently of the paired LLM