Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Deepgram Nova-3 · Evaluations · Kaino
D
model evaluation

Deepgram Nova-3

Deepgram

High-accuracy Deepgram speech-to-text model for batch and streaming transcription, with multilingual support and self-serve terminology customization.

modelsource:deepgram.comspeech-to-texttranscriptionstreamingbatchmultilingualDeepgram
72.1KAINO SCORERecommended
Evaluated Jul 31, 202612 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Speed & availability91
  • Technical capability87
  • Developer experience84
  • Multimodal & I/O81
  • Risk & evidence78
  • Cost effectiveness74
  • Adoption signal74
  • Pricing clarity70
  • Coding & agentic42
  • Reasoning & knowledge40

Kainotomic evaluation

Nova-3 is a specialized ASR model rather than a general LLM. Artificial Analysis reports 5.2% AA-WER for non-streaming use, 320.8× median speed, and strong realtime latency: 0.057 seconds to first partial and 0.06 seconds to final transcription. Official materials support batch and streaming APIs, multilingual use, and terminology customization. Its technical and audio-I/O scores therefore exceed broad LLM anchors on transcription-specific throughput, while remaining narrower than Gemini 3.5 Flash’s broader multimodal capability. Coding, agentic work, and general reasoning score well below Hermes, Grok 3, Kimi K2.5, and Claude Opus 4.8 because Nova-3 is not positioned or evaluated as a code-generating or instruction-following model. DeepSWE, SWE-bench, LiveCodeBench, Terminal-Bench/Aider, LMArena, and Arena-Hard provide no comparable model score. The Deepgram CLI is useful developer tooling, but its Aider detection does not establish Nova-3 coding-agent performance. At $4.30 per 1,000 audio minutes in Artificial Analysis’ non-streaming comparison, cost effectiveness is solid rather than leading without direct like-for-like price evidence. Published third-party WER and latency figures support a high speed score; official documentation and CLI support a strong developer-experience score. Pricing clarity is moderated because the supplied official documentation does not substantiate a complete pricing schedule. Evidence is comparatively strong for ASR, but weak or inapplicable for general intelligence claims.

Strengths

  • Independent Artificial Analysis results show 5.2% AA-WER, 320.8× median non-streaming speed, and 0.057-second time to first partial transcript.
  • Official support for both batch and realtime transcription, multilingual operation, and self-serve terminology customization.
  • Documented API and terminal CLI provide practical integration paths.

Caveats

  • Nova-3 is an ASR model; its scores for coding, agents, and general reasoning are intentionally low and are not measures of transcription quality.
  • The $4.30 per 1,000 audio minutes figure is from Artificial Analysis’ non-streaming comparison, not a complete official pricing analysis.
  • Audio transcription capability is narrower than the multimodal scope of general-purpose models such as Gemini 3.5 Flash.