Audio & Speech
This category covers models that generate or transcribe audio — text-to-speech and voice-cloning models, speech-to-text/transcription systems, and generative music/audio tools. It’s a more mature, commercially deployed category than image or video generation, since voice AI (call-center automation, voice assistants, dubbing) has clearer, already-proven enterprise use cases. GROUNDING tracks new speech/audio-model releases and the latency/naturalness benchmarks that determine whether a voice model is usable in a real-time conversational product.
At a glance
- 2 tracked models
- Most recently updated: Audio Flamingo 3 (2026-06-18)
Most actively covered
FAQ
What is the Audio & Speech category?
Text-to-speech, speech-to-text, and audio-generation models.
How many models are in this category?
GROUNDING currently tracks 2 models in Audio & Speech.
How current is this hub?
The most recently updated entry is Audio Flamingo 3, last mentioned 2026-06-18; the radar refreshes hourly.