Skip to content

Audio & Speech

This category covers models that generate or transcribe audio — text-to-speech and voice-cloning models, speech-to-text/transcription systems, and generative music/audio tools. It’s a more mature, commercially deployed category than image or video generation, since voice AI (call-center automation, voice assistants, dubbing) has clearer, already-proven enterprise use cases. GROUNDING tracks new speech/audio-model releases and the latency/naturalness benchmarks that determine whether a voice model is usable in a real-time conversational product.

At a glance

Most actively covered

FAQ

What is the Audio & Speech category?

Text-to-speech, speech-to-text, and audio-generation models.

How many models are in this category?

GROUNDING currently tracks 2 models in Audio & Speech.

How current is this hub?

The most recently updated entry is Audio Flamingo 3, last mentioned 2026-06-18; the radar refreshes hourly.

2 pages with this tag.