arrow
Return

Dual transform based joint learning single channel speech separation using generative joint dictionary learning

delete2022-04-02
delete2
PRE
AI
M
Md. Imran Hossain
T
Tarek Hasan Al Mahmud
M
Md Shohidul Islam
M
Md. Bipul Hossen
R
Rashid Khan
Z
Zhongfu Ye *
DOI:10.1007/s11042-022-12816-0delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Single channel speech separation (SS) is highly significant in many real-world speech processing applications such as hearing aids, automatic speech recognition, control humanoid robots, and cocktail-party issues. The performance of the SS is crucial for these applications, but better accuracy has yet to be developed. Some researchers have tried to separate speech using only the magnitude part, and some are tried to solve complex domains. We propose a dual transform SS method that serially uses the dual-tree complex wavelet transform (DTCWT) and short-term Fourier transform (STFT), and jointly learns the magnitude, real and imaginary parts of the signal applying a generative joint dictionary learning (GJDL). At first, the time-domain speech signal is decomposed by DTCWT, which produces a set of subband signals. Then STFT is connected to each subband signal, which converts each subband signal to the time-frequency domain and builds a complex spectrogram that prepares three parts like real, imaginary and magnitude for each subband signal. Next, we utilize the GJDL approach for making the joint dictionaries, and then the batch least angle regression with a coherence criterion (LARC) algorithm is used for sparse coding. Afterward, computes the initially estimated signals in two different ways, one by considering only the magnitude part and another by considering real and imaginary components. Finally, we apply the Gini index (GI) to the initially estimated signals to achieve better accuracy. The proposed algorithm demonstrates the best performance in all considered evaluation metrics compared to the mentioned algorithms.
Keywords:
Speech separation (SS)
Dual-tree complex wavelet transform (DTCWT)
Generative joint dictionary learning (GJDL)
Short-time Fourier transform (STFT)
Gini index (GI)

Journal

Multimedia Tools and Applications cover
Multimedia Tools and Applications
IF:
3
Papers:
2.0W
Citations:
3.2W

Organization

U
university of science & technology of china, cas
Scholars:
3.2W
Papers: 2.7W
Citations: 74
I
Islamic University
Scholars:
688
Papers: 350
Citations: 392
C
chinese academy of sciences
Scholars:
56.7W
Papers: 45.0W
Citations: 704
researcher View more organizations
Cited Papers

Cited Papers

Spatiotemporal Attention Enhances Lidar-Based Robot Navigation in Dynamic Environments
err2024-05-01
err0
errOAAI
errJorge de Heuvel; Xiangyu Zeng; Weixian Shi; Tharun Sethuraman; Maren Bennewitz
errShare
errSave
Sprint Ability: How Well Does Your Software Exploit Bursts in Processing Capacity?
err2016-07-01
err0
PREAI
errNathaniel Morris; Siva Meenakshi Renganathan; Christopher Stewart; Robert Birke; Lydia Chen
errShare
errSave
Single Channel multi-speaker speech Separation based on quantized ratio mask and residual network
err2020-08-26
err3
PREAI
errKe, Shanfa; Hu, Ruimin; Wang, Xiaochen; Wu, Tingzhao; Li, Gang; Wang, Zhongyuan
errShare
errSave
The distonic HC+ (OH) OĊH2 radical cation: a stable isomer of ionized methyl formate
err1991-11-01
err0
PREAI
errR. Flammang; M. Plisnier; G. Leroy; M. Sana; Minh Tho Nguyen; L.G. Vanquickenborne
errShare
errSave
Design of High-Performance Microprocessor Circuits
err
IF0
err2000-01-01
err0
PREAI
errAnantha Chandrakasan; William J. Bowhill; Frank Fox
errShare
errSave
researcher View more