Rethage D, Pons J, Serra X. A Wavenet for Speech Denoising. 43rd IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP2018)

Thesis linked to the implementation of the María de Maeztu Strategic Research Program.

Open access to PhD thesis carried out at the Department can be found at TDX

Please visit these pages for information on our PhD, MSc and BSc programs.

Back Rethage D, Pons J, Serra X. A Wavenet for Speech Denoising. 43rd IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP2018)

Rethage D, Pons J, Serra X. A Wavenet for Speech Denoising. 43rd IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP2018)

Currently, most speech processing techniques use magnitude spectrograms as front-end and are therefore by default discarding part of the signal: the phase. In order to overcome this limitation, we propose an end-to-end learning method for speech denoising based on Wavenet. The proposed model adaptation retains Wavenet's powerful acoustic modeling capabilities, while significantly reducing its time-complexity by eliminating its autoregressive nature. Specifically, the model makes use of non-causal, dilated convolutions and predicts target fields instead of a single target sample. The discriminative adaptation of the model we propose, learns in a supervised fashion via minimizing a regression loss. These modifications make the model highly parallelizable during both training and inference. Both computational and perceptual evaluations indicate that the proposed method is preferred to Wiener filtering, a common method based on processing the magnitude spectrogram.

Additional material

Link: http://mtg.upf.edu/node/3800

DTIC MdM Strategic Program: Artificial and Natural Intelligence for ICT and beyond

Rethage D, Pons J, Serra X. A Wavenet for Speech Denoising. 43rd IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP2018)

Related Assets