Zyphra
Open-source MoE text-to-speech model for real-time, high-fidelity voice cloning with Apache 2.0 weights and hosted or self-hosted inference paths.
ZONOS2 is a specialized TTS system rather than a general-purpose language model. Official materials and its technical report describe an 8B MoE architecture with 900M active parameters, streaming-oriented latency, real-time inference, voice cloning, Apache-2.0 weights, and hosted or self-hosted deployment. That makes its speech-output and deployment profile stronger for this narrow workload than text-centric anchors such as Gemma 3, while its overall technical breadth remains far below GPT-5.5, Claude Opus 4.8, Grok 3, and Kimi K2.5. The open weights, GitHub repository, Hugging Face distribution, and dual hosted/self-managed paths support above-average cost-effectiveness and developer experience for teams able to operate audio inference. The latency and throughput claims support a strong speed score, but are vendor/report claims rather than independently normalized serving measurements. Unlike Gemini 3.5 Flash, ZONOS2 does not offer demonstrated broad multimodal input/output coverage; its I/O score reflects audio generation and voice-reference use, not vision, general reasoning, or tool use. Coding and reasoning scores are intentionally low because ZONOS2 is not positioned or evidenced as a coding or general-reasoning model. DeepSWE and LiveCodeBench list no result; no SWE-bench, Terminal-Bench, Aider, LMArena, or Arena-Hard result was supplied. Published usage, pricing, safety controls, and independent quality comparisons are also limited in the supplied evidence. This supports lower public-signal, pricing-clarity, and evidence-quality scores than the established general-model anchors.