Definition·
Capabilities

Speech-to-Text

Technology that automatically transcribes spoken language into written text.

Detailed explanation

Speech-to-Text (STT) or ASR (Automatic Speech Recognition) relies on deep learning models that handle many languages, accents and background noise. It powers dictation, meeting transcription, captions and voice commands.

Examples

Zoom meeting transcription
YouTube auto-captions
iPhone voice dictation
Whisper (OpenAI)

Frequently asked questions

How accurate is it?

Top models exceed 95% in good conditions; noise, accents and domain jargon reduce quality.

Related terms

Last updated: 7/15/2026

Talent AI

Turn theory into practice

Post a mission or join the community of top AI, Data and Machine Learning experts.