arrow
返回

Bayesian Multichannel Audio Source Separation Based on Integrated Source and Spatial Models

delete2018-04-01
delete13
PRE
AI
K
Kousuke Itakura
Y
Yoshiaki Bando
E
Eita Nakamura
K
Katsutoshi Itoyama
K
Kazuyoshi Yoshii *
T
Tatsuya Kawahara
DOI:10.1109/TASLP.2017.2789320delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This paper presents new statistical methods of multichannel audio source separation based on unified source and spatial models that, respectively, represent the generative process of latent source spectrograms and that of observed mixture spectrograms. One possibility of the source model is a factor model based on nonnegative matrix factorization that represents each time-frequency (TF) bin as the weighted sum of basis spectra. Another possibility is a mixture model inspired by latent Dirichlet allocation that exclusively classifies each TF bin into one of basis spectra. Similarly, the spatial model can either be a factor model that represents each TF bin as the weighted sum of source spectra or a mixture model that classifies each bin into one of those spectra. To unify these models in a principled manner and incorporate prior knowledge of a microphone array, we propose hierarchical Bayesianmodels of all the source-spatial combinations (factor-factor, mixture-factor, factor-mixture, and mixture-mixture models) and derive efficient Gibbs sampling algorithms for posterior inference. Experimental results showed that the proposed unified models outperformed the state-of-the-art method using only the spatial mixture model. Among the four unified models, the spatial factor model tended to work better than the spatial mixture model in exchange for larger computational cost, and the choice of source models had a little impact on the performance and computational cost.
Keyword:
Multichannel source separation
latent Dirichlet allocation
nonnegative matrix factorization
Bayesian models
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
论文数:
2.6K
被引数:
1.1W

机构

K
Kyoto University
学者数:
5.1W
论文数: 4.6W
被引数: 6.1W
引用论文

引用论文

Sprint Ability: How Well Does Your Software Exploit Bursts in Processing Capacity?
err2016-07-01
err0
PREAI
errNathaniel Morris; Siva Meenakshi Renganathan; Christopher Stewart; Robert Birke; Lydia Chen
err分享
err收藏
Phase Equilibria in the Co-W-Zr Ternary System at 1200 and 1300 °C
err2019-11-29
err0
PREAI
errXingjun Liu; Chenbin Luo; Mujing Yang; Shuiyang Yang; Jinbin Zhang; Yixiong Huang; Jiajia Han; Yong Lu; Cuiping Wang
err分享
err收藏
err分享
err收藏
学者 查看更多内容