Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
FastSpeech is a neural text-to-speech model developed by Yi Ren of Zhejiang University. The model uses a feed-forward Transformer network to generate mel-spectrograms in parallel, addressing speed and robustness issues in autoregressive TTS. Experiments on the LJSpeech dataset show the model speeds up mel-spectrogram generation by 270x and end-to-end synthesis by 38x compared to an autoregressive Transformer baseline.
License is listed as Open Access (green); specific terms should be verified.