Deepgram
High-accuracy Deepgram speech-to-text model for batch and streaming transcription, with multilingual support and self-serve terminology customization.
Nova-3 is a specialized ASR model rather than a general LLM. Artificial Analysis reports 5.2% AA-WER for non-streaming use, 320.8× median speed, and strong realtime latency: 0.057 seconds to first partial and 0.06 seconds to final transcription. Official materials support batch and streaming APIs, multilingual use, and terminology customization. Its technical and audio-I/O scores therefore exceed broad LLM anchors on transcription-specific throughput, while remaining narrower than Gemini 3.5 Flash’s broader multimodal capability. Coding, agentic work, and general reasoning score well below Hermes, Grok 3, Kimi K2.5, and Claude Opus 4.8 because Nova-3 is not positioned or evaluated as a code-generating or instruction-following model. DeepSWE, SWE-bench, LiveCodeBench, Terminal-Bench/Aider, LMArena, and Arena-Hard provide no comparable model score. The Deepgram CLI is useful developer tooling, but its Aider detection does not establish Nova-3 coding-agent performance. At $4.30 per 1,000 audio minutes in Artificial Analysis’ non-streaming comparison, cost effectiveness is solid rather than leading without direct like-for-like price evidence. Published third-party WER and latency figures support a high speed score; official documentation and CLI support a strong developer-experience score. Pricing clarity is moderated because the supplied official documentation does not substantiate a complete pricing schedule. Evidence is comparatively strong for ASR, but weak or inapplicable for general intelligence claims.