MODELS
Higgs Audio v3 TTS
by Boson AI
Overview
4B conversational text-to-speech model for low-latency voice agents with multilingual speech, voice cloning, and inline emotion/style control.
Details
Higgs Audio v3 TTS is Boson AI’s 4B conversational text-to-speech model for expressive voice AI. Official Boson sources describe it as built for voice-chat and low-latency voice agents, with expressive conversational speech across 100+ languages, zero-shot/reference-audio voice cloning, streaming speech endpoint examples, and inline text controls for emotion, style, prosody, pauses, pitch, speed, expressiveness, and sound effects. The documented API model name is `higgs-audio-v3-tts`.
When to Use
Use for voice-agent or voice-chat experiences that need low-latency expressive conversational speech rather than plain read-aloud TTS. Use when a TTS workflow needs multilingual output reference-audio voice cloning and inline control tokens for emotion style pauses pitch speed prosody or sound effects. Evaluate when you want an officially documented API model with a public demo docs Hugging Face model page and GitHub project links.
Getting Started
- Read the official overview at https://docs.boson.ai/models/higgs-audio-tts/overview and note the API model name `higgs-audio-v3-tts`.
- Try the official demo at https://www.boson.ai/demo/tts to assess speech quality and controllability.
- Review Boson’s tags documentation at https://docs.boson.ai/models/higgs-audio-tts/tags for supported inline emotion
- style
- sound-effect
- speed
- pause
- pitch
- and expressiveness controls.
- Inspect the GitHub repository at https://github.com/boson-ai/higgs-audio or the Hugging Face page at https://huggingface.co/bosonai/higgs-audio-v3-tts-4b before production evaluation.
Key Features
- •4B conversational text-to-speech model for expressive real-speech style output.
- •Supports expressive conversational speech across 100+ languages
- •according to Boson’s announcement sources.
- •Zero-shot/reference-audio voice cloning is described in official Boson docs and launch materials.
- •Inline text tags can control emotion
- •style
- •vocalized sound effects
- •speed
- •pauses
- •pitch
- •prosody
- •and expressiveness.
- •Streaming speech endpoint examples are documented for the API model `higgs-audio-v3-tts`.
Capabilities
- •text-to-speech
- •streaming speech generation
- •multilingual speech synthesis
- •voice cloning
- •emotion and style control
- •prosody control
- •voice-agent audio generation
Last updated Jul 31, 2026