arrow
Return

Training-Based Multiple Source Tracking Using Manifold-Learning and Recursive Expectation-Maximization

delete2023-01-01
delete0
PRE
AI
A
Avital Bross
S
Sharon Gannot *
DOI:10.1109/TASLP.2023.3245414delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In this paper we propose a data-driven approach for multiple speaker tracking in reverberant enclosures. The speakers are uttering, possibly overlapping, speech signals while moving in the environment. The method comprises two stages. The first stage executes a single source localization using semi-supervised learning on multiple manifolds. The second stage, which is unsupervised, uses time-varying maximum likelihood estimation for tracking. The feature vectors, used by both stages, are the relative transfer functions (RTFs), which are known to be related to source positions. The number of sources is assumed to be known while the microphone positions are unknown. In the training stage, a large database of RTFs is given. A small percentage of the data is attributed with exact positions (namely, labelled data) and the rest is assumed to be unlabelled, i.e. the respective position is unknown. Then, a nonlinear, manifold-based, mapping function between the RTFs and the source positions is inferred. Applying this mapping function to all unlabelled RTFs constructs a dense grid of localized sources. In the test phase, this RTF grid serves as the centroids for a Mixture of Gaussians (MoG) model. The MoG parameters are estimated by applying a recursive variant of the expectation-maximization (EM) procedure that relies on the sparsity and intermittency of the speech signals. We present a comprehensive simulation study in various reverberation levels, including static and dynamic scenarios, for both two or three (partially) overlapping speakers. For the dynamic case we provide simulations with several speakers trajectories, including intersecting sources. The proposed scheme outperforms baseline methods that use a simpler propagation model in terms of localization accuracy and tracking capabilities.
Keywords:
Location awareness
Microphones
Acoustics
Speech processing
Hidden Markov models
Manifolds
Feature extraction
Manifold learning
multiple source tracking
recursive expectation-maximization
speech sparsity

Journal

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
Papers:
2.6K
Citations:
1.1W

Organization

B
Bar Ilan University
Scholars:
9.7K
Papers: 8.5K
Citations: 59
Cited Papers

Cited Papers

Attention Deficit Hyperactivity Disorder (ADHD) is Associated with Early Onset Substance Use Disorders
err1997-08-01
err0
PREAI
errTIMOTHY E. WILENS; JOSEPH BIEDERMAN; ERIC MICK; STEPHEN V. FARAONE; THOMAS SPENCER
errShare
errSave
errShare
errSave
errShare
errSave
errShare
errSave
Pesticide use and fatal injury among farmers in the Agricultural Health Study
err2012-03-15
err0
errOAAI
errJenna K. Waggoner; Paul K. Henneberger; Greg J. Kullman; David M. Umbach; Freya Kamel; Laura E. Beane Freeman; Michael C. R. Alavanja; Dale P. Sandler; Jane A. Hoppin
errShare
errSave
Low oxygen levels with high redox heterogeneity in the late Ediacaran shallow ocean: Constraints from I/(Ca + Mg) and Ce/Ce* of the Dengying Formation, South China
err2022-08-09
err0
PREAI
errYi Ding; Wei Sun; Shugen Liu; Jirong Xie; Dongjie Tang; Xiqiang Zhou; Limin Zhou; Zhiwu Li; Jinmin Song; Zeqi Li; Hongyuan Xu; Pan Tang; Kang Liu; Wenjun Li; Daizhao Chen
errShare
errSave
errShare
errSave
Raking the Cocktail Party
err2015-08-01
err39
errOAAI
errDokmanic, Ivan; Scheibler, Robin; Vetterli, Martin
errShare
errSave
researcher View more