Home Knowledge Base Role in TTS pipeline

Neural vocoders convert acoustic features (mel spectrograms) back into high-fidelity audio waveforms. Role in TTS pipeline: Text leads to acoustic model leads to mel spectrogram leads to vocoder leads to audio waveform. Vocoder is final synthesis stage. Why needed: Mel spectrograms are compact representation, but contain no phase information needed for waveform. Vocoder reconstructs plausible phase and generates samples. Key architectures: Autoregressive: WaveNet (slow, high quality, sample-by-sample), WaveRNN. Non-autoregressive: HiFi-GAN (fast, excellent quality), UnivNet, Vocos. GAN vocoders: Generator produces waveform, discriminators judge quality. Multi-scale and multi-period discriminators. Training: Reconstruct original audio from mel spectrogram, GAN loss + feature matching + mel reconstruction. Quality vs speed: WaveNet: 1000x slower than real-time. HiFi-GAN: 1000x faster than real-time, comparable quality. Universal vocoders: Work across speakers/conditions vs speaker-specific. Integration: End-to-end models (VITS) combine acoustic model and vocoder. HiFi-GAN made high-quality neural TTS practical.

neural vocoderaudio

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.