Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Sonic 3.5 · Evaluations · Kaino
S
model evaluation

Sonic 3.5

Cartesia

Sonic 3.5 is Cartesia’s fast, natural streaming TTS model with sub-90ms latency, 42-language support, and stable dated snapshots.

modelmultilingualcartesia
59.6KAINO SCORENot recommended
Evaluated Jul 31, 202612 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Speed & availability94
  • Technical capability84
  • Developer experience78
  • Risk & evidence76
  • Multimodal & I/O72
  • Adoption signal72
  • Cost effectiveness62
  • Pricing clarity45
  • Reasoning & knowledge8
  • Coding & agentic5

Kainotomic evaluation

Sonic 3.5 is a specialized streaming TTS model, not a general-purpose language or coding system. Official documentation supports 42-language speech generation, dated production snapshots, and sub-90 ms latency; Cartesia’s H100 result reports 50 ms P50 time-to-first-audio and 105 characters/s at concurrency one. Artificial Analysis reports a 1,207 Speech Arena Elo (±14, 2,085 appearances), placing it near the leading quality tier, though behind Fun-Realtime-TTS’s cited 1,219 result. Its technical and audio I/O scores are therefore stronger than broad models such as Gemma 3 or Grok 3 only within realtime speech synthesis, while its coding and reasoning scores are far below those general models and especially Claude Opus 4.8, GPT-5.5, and Kimi K2.5. The speed score exceeds the published general-model anchors because the evidence is direct TTS latency/throughput evidence, not token-generation inference. Stable model snapshots and provider documentation support a solid developer-experience score. Cost effectiveness is only provisional: Artificial Analysis evaluates price per million characters, but the supplied evidence provides no price figure or plan detail. Pricing clarity is consequently weak. Speech Arena provides meaningful preference evidence, but it is not LMArena/Arena-Hard and does not establish text reasoning quality. No Sonic 3.5 results appear in DeepSWE or LiveCodeBench; SWE-bench evidence was not found, and the supplied Terminal-Bench/Aider material does not provide a result.

Strengths

  • Direct evidence of very low TTS latency and strong streaming throughput
  • 42-language support and production-oriented dated snapshots
  • Strong public TTS preference signal: 1,207 Speech Arena Elo across 2,085 appearances

Caveats

  • Scores for coding and reasoning reflect an out-of-scope TTS model, not poor performance on a claimed capability
  • Speech Arena quality evidence is TTS-specific and should not be compared directly with LLM arena scores
  • No supplied price amount, licensing terms, or service-level availability commitment