Higgs Audio v3 TTS logo

MODELS

Higgs Audio v3 TTS

by Boson AI

modelsource:boson.aitext-to-speechttsvoice-aivoice-agentsmultilingualvoice-cloningstreamingemotion-controlprosody-controlBoson AI

Overview

4B conversational text-to-speech model for low-latency voice agents with multilingual speech, voice cloning, and inline emotion/style control.

Details

Higgs Audio v3 TTS is Boson AI’s 4B conversational text-to-speech model for expressive voice AI. Official Boson sources describe it as built for voice-chat and low-latency voice agents, with expressive conversational speech across 100+ languages, zero-shot/reference-audio voice cloning, streaming speech endpoint examples, and inline text controls for emotion, style, prosody, pauses, pitch, speed, expressiveness, and sound effects. The documented API model name is `higgs-audio-v3-tts`.

When to Use

Use for voice-agent or voice-chat experiences that need low-latency expressive conversational speech rather than plain read-aloud TTS. Use when a TTS workflow needs multilingual output reference-audio voice cloning and inline control tokens for emotion style pauses pitch speed prosody or sound effects. Evaluate when you want an officially documented API model with a public demo docs Hugging Face model page and GitHub project links.

Getting Started

  1. Read the official overview at https://docs.boson.ai/models/higgs-audio-tts/overview and note the API model name `higgs-audio-v3-tts`.
  2. Try the official demo at https://www.boson.ai/demo/tts to assess speech quality and controllability.
  3. Review Boson’s tags documentation at https://docs.boson.ai/models/higgs-audio-tts/tags for supported inline emotion
  4. style
  5. sound-effect
  6. speed
  7. pause
  8. pitch
  9. and expressiveness controls.
  10. Inspect the GitHub repository at https://github.com/boson-ai/higgs-audio or the Hugging Face page at https://huggingface.co/bosonai/higgs-audio-v3-tts-4b before production evaluation.

Key Features

  • 4B conversational text-to-speech model for expressive real-speech style output.
  • Supports expressive conversational speech across 100+ languages
  • according to Boson’s announcement sources.
  • Zero-shot/reference-audio voice cloning is described in official Boson docs and launch materials.
  • Inline text tags can control emotion
  • style
  • vocalized sound effects
  • speed
  • pauses
  • pitch
  • prosody
  • and expressiveness.
  • Streaming speech endpoint examples are documented for the API model `higgs-audio-v3-tts`.

Capabilities

  • text-to-speech
  • streaming speech generation
  • multilingual speech synthesis
  • voice cloning
  • emotion and style control
  • prosody control
  • voice-agent audio generation

Last updated Jul 31, 2026