arrow
返回

Efficient Personalized Speech Enhancement Through Self-Supervised Learning

delete2022-10-01
delete7
delete
OA
AI
A
Aswin Sivaraman
M
Minje Kim *
DOI:10.1109/JSTSP.2022.3181782delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This work presents self-supervised learning methods for monaural speaker-specific (i.e., personalized) speech enhancement models. While general-purpose models must broadly address many speakers, personalized models can adapt to a particular speaker's voice, expecting to solve a narrower problem. Hence, personalization can achieve more optimal performance in addition to reducing computational complexity. However, naive personalization methods can inconveniently require clean speech from the target user, e.g., due to subpar recording conditions. To this end, we pose personalization as either a zero-shot task, in which no clean speech of the target speaker is used, or a few-shot learning task, which is to minimize the duration of the clean speech used for transfer learning. With this paper, we propose self-supervised learning methods as a solution to both zero- and few-shot personalization tasks. The proposed methods learn the personalized speech features from unlabeled data (i.e., in-the-wild noisy recordings from the target user) rather than from the clean sources. We investigate three different self-supervised learning mechanisms. We set up a pseudo speech enhancement problem as a pretext task, which pretrains the models to estimate noisy speech as if it were the clean target. Contrastive learning and data purification methods regularize the loss function of the pseudo enhancement problem, overcoming the limitations of learning from unlabeled data. We assess our methods by personalizing the well-known ConvTasNet architecture to twenty different target speakers. The results show that self-supervision-based personalization improves the original ConvTasNet's enhancement quality with fewer model parameters and less clean data from the target user.
Keyword:
Speech enhancement
Noise measurement
Data models
Training
Task analysis
Adaptation models
Recording
Data efficiency
model complexity
personalized speech enhancement
self-supervised learning

期刊

IEEE Journal of Selected Topics in Signal Processing 封面图
IEEE Journal of Selected Topics in Signal Processing
IF:
13.7
论文数:
1.9K
被引数:
1.1W

机构

I
indiana university system
学者数:
4.0W
论文数: 3.5W
被引数: 38
I
Indiana University Bloomington
学者数:
1.9W
论文数: 1.5W
被引数: 2.8W
引用论文

引用论文

err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Sprint Ability: How Well Does Your Software Exploit Bursts in Processing Capacity?
err2016-07-01
err0
PREAI
errNathaniel Morris; Siva Meenakshi Renganathan; Christopher Stewart; Robert Birke; Lydia Chen
err分享
err收藏
学者 查看更多内容