arrow
Return

Multimodal System for Audio Scene Source Counting and Analysis

delete2022-01-01
delete2
delete
OA
AI
M
Michael Nigro *
S
Sridhar Krishnan
DOI:10.1109/TASLP.2022.3156795delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Audio scene analysis (ASA) is a challenging and multifaceted task in audio signal processing that uncovers information about the nature of an audio recording. Regardless of the analysis goal, a number of audio sources are observed in any audio scene. However, this consideration is usually not explored or given considerable thought in research. This work aims to demonstrate the utility of audio source counting with a novel solution consisting of a multimodal system for ASA. Both speaker counting and sound event counting techniques use deep neural networks (DNN) to predict the number of sources. We are able to present competitive results for audio source counting by achieving prediction accuracy of 46.03% and 89.57% with a margin of error of +/- 1 for speaker counting, which outperforms state-of-the-art systems for similar tasks. For sound event counting we achieve 50.55% and 86.59% prediction accuracy and accuracy with a margin of error of +/- 1, respectively, that establishes a clear baseline. Our system also demonstrates real-time aspects with an overall processing time of similar to 0.4614 s per audio recording.
Keywords:
Audio scene analysis
source counting
speaker count estimation

Journal

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
Papers:
2.6K
Citations:
1.1W

Organization

T
Toronto Metropolitan University
Scholars:
6.0K
Papers: 7.0K
Citations: 6.4K