Home Knowledge Base Automatic Speech Recognition (ASR)

Automatic Speech Recognition (ASR) is the deep learning system that converts spoken audio into text — processing raw audio waveforms through neural encoder-decoder architectures that learn to map acoustic features to linguistic tokens, achieving human-level transcription accuracy across languages and accents through end-to-end training on hundreds of thousands of hours of paired audio-text data.

Architecture Evolution

End-to-End Architectures

Whisper (OpenAI, 2022)

Trained on 680,000 hours of weakly-supervised web audio in 99 languages. Encoder-decoder Transformer with multitask training: transcription, translation, language identification, timestamp prediction — all controlled by text prompts. Achieves near-human accuracy on English without any fine-tuning. Demonstrated that scaling data (not architecture novelty) was the primary bottleneck for robust ASR.

Audio Feature Processing

Automatic Speech Recognition is the interface between human speech and machine understanding — a technology that has progressed from 50% word error rates to human-parity accuracy in a decade, enabling voice assistants, real-time captioning, and multilingual communication at planetary scale.

speech recognition asrwhisper speech modelconnectionist temporal classification ctcend to end speechautomatic speech recognition

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.