arrow
返回

Differentiable Measures for Speech Spectral Modeling

delete2022-01-01
delete2
delete
OA
AI
M
Miguel Arjona Ramírez *
W
Wesley Beccaro
D
Demóstenes Zegarra Rodríguez
R
Renata Lopes Rosa
DOI:10.1109/ACCESS.2022.3150728delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Autoregressive models for the envelope of speech power spectral densities (PSDs) are refined by the self-supervised spectral learning machine (S3LM) provided with differentiable spectral objective functions, including the Itakura-Saito divergence (ISD), the Kullback-Leibler divergence (KLD), the reverse KLD (RKLD) and the log spectral distortion (LSD), which display more significant results. However, in order to assess the models more perceptually, a method is proposed based upon perturbations around perfect reconstruction analysis-synthesis configurations. In the cross-excitation analysis-synthesis assessment (CEASA) method, the residual signals generated by analysis filters of the spectral models are injected as excitation into the synthesis filters derived from the same and other models in order to be evaluated by the perceptual evaluation of speech quality (PESQ) and Itakura divergence (ID), which are averaged over a set of models obtained using the objective functions mentioned above. The results lead to a superior performance when the RKLD is used as the loss function for the estimation of the spectral models with the ISD ranking close behind. The focus of these divergences on the spectral peaks is argued and pointed as the most important factor for this behavior. Specifically, using the PESQ scores obtained with CEASA, the RKLD loss is found to improve the performance by 1.0%, 4.0% and 19.3% with respect to the open-loop analysis, the KLD and the LSD models, respectively, while the corresponding improvements for the ISD loss are 0.1%, 3.0% and 18.2%, and the RKLD models excel the ISD models by 1.0% on average. Even though the spectral measures alone are not able to unequivocally distinguish the better of the two, CEASA is shown to have enough sensitivity to distinguish their performances. In summary, the learning machine S3LM fits models for the short-term spectral envelope of speech and, for the evaluation of its performance under several differentiable loss functions, the CEASA assessment tool has been developed. In addition, CEASA may be used for other assessments connected with speech analysis and synthesis.
Keyword:
Analytical models
Power harmonic filters
Modeling
Predictive models
Loss measurement
Electronic mail
Autocorrelation
Autoregressive processes
machine learning algorithms
prediction methods
self-supervised learning
speech analysis
spectral analysis

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

Universidade Federal de Lavras 封面图
Universidade Federal de Lavras
学者数:
5.7K
论文数: 3.3K
被引数: 3.5K
U
universidade de sao paulo
学者数:
10.6W
论文数: 6.7W
被引数: 93
引用论文

引用论文

The continual impact of the Paris System on urine cytology, a 3‐year experience
err2019-12-08
err0
PREAI
errNicholas Stanzione; Tagreed Ahmed; Po Chu Fung; Diancai Cai; David Y. Lu; Lauren C. Sumida; Neda A. Moatamed
err分享
err收藏
Full-Band LPCNet: A Real-Time Neural Vocoder for 48 kHz Audio With a CPU
err2021-01-01
err9
PREAI
errMatsubara, Keisuke; Okamoto, Takuma; Takashima, Ryoichi; Takiguchi, Tetsuya; Toda, Tomoki; Shiga, Yoshinori; Kawai, Hisashi
err分享
err收藏
Sulphasomizole (5-p-Aminobenzenesulphonamido-3-Methylisothiazole): A New Antibacterial Sulphonamide
err1960-04-01
err0
PREAI
errA. ADAMS; W. A. FREEMAN; A. HOLLAND; D. HOSSACK; J. INGLIS; J. PARKINSON; H. W. READING; K. RIVETT; R. SLACK; R. SUTHERLAND; R. WIEN
err分享
err收藏
Front cover封面
err2013-01-01
err0
PREAI
err
err分享
err收藏
Transient Cortical Blindness Following Bypass Graft Angiography
err1995-10-01
err0
PREAI
errJunya Kamata; Kenichi Fukami; Hiroaki Yoshida; Yoshimi Mizunuma; Naoki Moriai; Toshitake Takino; Shunichi Hosokawa; Koya Hashimoto; Kenji Nakai; Kohei Kawazoe; Katsuhiko Hiramori
err分享
err收藏
学者 查看更多内容