Return
Audio Super-Resolution With Robust Speech Representation Learning of Masked Autoencoder
S
S
H
S
DOI:10.1109/TASLP.2023.3349053.png)
Abstract
En 中文
This paper proposes Fre-Painter, a high-fidelity audio super-resolution system that utilizes robust speech representation learning with various masking strategies. Recently, masked autoencoders have been found to be beneficial in learning robust representations of audio for speech classification tasks. Following these studies, we leverage these representations and investigate several masking strategies for neural audio super-resolution. In this paper, we propose an upper-band masking strategy with the initialization of the mask token, which is simple but efficient for audio super-resolution. Furthermore, we propose a mix-ratio masking strategy that makes the model robust for input speech with various sampling rates. For practical applicability, we extend Fre-Painter to a text-to-speech system, which synthesizes high-resolution speech using low-resolution speech data. The experimental results demonstrate that Fre-Painter outperforms other neural audio super-resolution models.
Keywords:
Superresolution
Task analysis
Speech processing
Self-supervised learning
Computational modeling
Decoding
Training
Audio super-resolution
bandwidth extension
self-supervised learning
masked autoencoder
audio synthesis
Journal
I
IF:
5.1
Papers:
2.6K
Citations:
1.1W
