arrow
返回

Dual transform based joint learning single channel speech separation using generative joint dictionary learning

delete2022-04-02
delete2
PRE
AI
M
Md. Imran Hossain
T
Tarek Hasan Al Mahmud
M
Md Shohidul Islam
M
Md. Bipul Hossen
R
Rashid Khan
Z
Zhongfu Ye *
DOI:10.1007/s11042-022-12816-0delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Single channel speech separation (SS) is highly significant in many real-world speech processing applications such as hearing aids, automatic speech recognition, control humanoid robots, and cocktail-party issues. The performance of the SS is crucial for these applications, but better accuracy has yet to be developed. Some researchers have tried to separate speech using only the magnitude part, and some are tried to solve complex domains. We propose a dual transform SS method that serially uses the dual-tree complex wavelet transform (DTCWT) and short-term Fourier transform (STFT), and jointly learns the magnitude, real and imaginary parts of the signal applying a generative joint dictionary learning (GJDL). At first, the time-domain speech signal is decomposed by DTCWT, which produces a set of subband signals. Then STFT is connected to each subband signal, which converts each subband signal to the time-frequency domain and builds a complex spectrogram that prepares three parts like real, imaginary and magnitude for each subband signal. Next, we utilize the GJDL approach for making the joint dictionaries, and then the batch least angle regression with a coherence criterion (LARC) algorithm is used for sparse coding. Afterward, computes the initially estimated signals in two different ways, one by considering only the magnitude part and another by considering real and imaginary components. Finally, we apply the Gini index (GI) to the initially estimated signals to achieve better accuracy. The proposed algorithm demonstrates the best performance in all considered evaluation metrics compared to the mentioned algorithms.
Keyword:
Speech separation (SS)
Dual-tree complex wavelet transform (DTCWT)
Generative joint dictionary learning (GJDL)
Short-time Fourier transform (STFT)
Gini index (GI)

期刊

Multimedia Tools and Applications 封面图
Multimedia Tools and Applications
IF:
3
论文数:
2.0W
被引数:
3.2W

机构

U
university of science & technology of china, cas
学者数:
3.2W
论文数: 2.7W
被引数: 74
I
Islamic University
学者数:
688
论文数: 350
被引数: 392
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
学者 查看更多机构
引用论文

引用论文

Sprint Ability: How Well Does Your Software Exploit Bursts in Processing Capacity?
err2016-07-01
err0
PREAI
errNathaniel Morris; Siva Meenakshi Renganathan; Christopher Stewart; Robert Birke; Lydia Chen
err分享
err收藏
Single Channel multi-speaker speech Separation based on quantized ratio mask and residual network
err2020-08-26
err3
PREAI
errKe, Shanfa; Hu, Ruimin; Wang, Xiaochen; Wu, Tingzhao; Li, Gang; Wang, Zhongyuan
err分享
err收藏
The distonic HC+ (OH) OĊH2 radical cation: a stable isomer of ionized methyl formate
err1991-11-01
err0
PREAI
errR. Flammang; M. Plisnier; G. Leroy; M. Sana; Minh Tho Nguyen; L.G. Vanquickenborne
err分享
err收藏
Design of High-Performance Microprocessor Circuits
err
IF0
err2000-01-01
err0
PREAI
errAnantha Chandrakasan; William J. Bowhill; Frank Fox
err分享
err收藏
学者 查看更多内容