Detailed explanation
Text-to-Speech (TTS) generates natural speech, sometimes cloned from a real voice. Recent models handle prosody, emotions, multiple languages and accents, enabling voice assistants, audiobooks and automatic dubbing.
Examples
Siri or Google Assistant voice
ElevenLabs to clone a voice
Audio reading of an article
Automatic video dubbing
Frequently asked questions
Fraud risks?
Voice cloning enables CEO scams; add extra identity checks for sensitive flows.
Related terms
Generative AI
A family of AI models that can create new content (text, image, audio, video, code) from a prompt.
API
An interface that lets an application call a service's functions — including an AI model — via HTTP requests.
Multimodal
The ability of an AI model to understand and generate multiple modalities (text, image, audio, video) inside a single system.
Speech-to-Text
Technology that automatically transcribes spoken language into written text.
Last updated: 7/15/2026