Fish Audio
Expressive text-to-speech model with natural-language emotion and prosody control, available through Fish Audio app/API and an open-source release.
Fish Audio S2 Pro is a specialized expressive TTS system rather than a general language or agent model. Official documentation, launch material, and the Fish Speech repository support API/app access, open-source availability, and natural-language emotion and prosody control. Artificial Analysis reports a 1122.82 TTS Quality Elo and evaluates price and generation speed, supporting a solid task-specific technical score, but the supplied evidence does not provide the underlying price or throughput values. Its audio-output focus gives it meaningful I/O capability, not broad multimodal understanding. Against the published anchors, S2 Pro is materially below GPT-5.5, Claude Opus 4.8, Kimi K2.5, and Gemma 3 for coding, agentic work, and general reasoning because it is not evidenced as a code or general-purpose LLM. It is more directly capable than those models for its narrow speech-synthesis use case, but that does not justify matching their broad technical-capability scores. Documentation plus an open repository make developer access stronger than a closed, undocumented TTS endpoint, though less mature than the SDK and tooling ecosystems reflected in the leading general-model anchors. Public signal is moderate: Voice Arena places the real-time variant 11th, at 963 Elo from 5,132 comparisons. This is useful preference evidence but is not directly interchangeable with the cited Artificial Analysis quality result. No SWE-bench result was found; DeepSWE, LiveCodeBench, and Terminal-Bench list no S2 Pro entry. Pricing and license terms were not supplied in sufficient detail, limiting cost, clarity, and evidence-confidence scores.