arrow
返回

Deep learning-based method for multiple sound source localization with high resolution and accuracy

delete2021-12-01
delete42
PRE
AI
S
Soo Young Lee
C
Chang, Jiho
S
Seung‐Chul Lee *
DOI:10.1016/j.ymssp.2021.107959delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Deep learning-based methods are attracting interest in sound source localization, showing promising results compared to conventional model-based approaches. While these deep learning-based methods have been mainly developed into two approaches, i.e., grid-based and grid-free methods, they inherently involve several limitations that the sound sources should be assumed on the grid points or the number of sound sources should be predefined when constructing a deep neural network's architecture. Breaking away from the existing methods' limitations, we propose a deep learning approach to fulfill multiple sound source localization with high resolution and accuracy, for whether the sound sources are located on the grid points or not. We first suggest a target function to obtain spatial source distribution maps, that can represent multiple sources' positional and strength information, even when the sources are placed off the grid points. While the multiple sound source localization is expanded by the proposed source map into image-to-image pixel-level prediction task, we then propose a fully convolutional neural network (FCN) with an encoder-decoder structure to estimate the multiple sources' positions and strength precisely. Based on the dataset acquired by one to three monopole sources on a square plane of 2.68 x 2.68 m, with a spiral array of 60 microphones at 1, 2, and 10 kHz, we assess both quantitative and qualitative results of the proposed model and demonstrate that our proposed model can achieve highly precise localization results regardless of frequency and the number of sound sources. Besides, we validate that high-resolution source distribution maps can be obtained by the proposed model, from which the positions and the strengths of sound sources are accurately predicted. Lastly, we compare the proposed model with several deconvolution methods, and the results show that the proposed deep learning model significantly outperforms the model-based methods. (C) 2021 Elsevier Ltd. All rights reserved.
Keyword:
Sound source localization
High resolution source map
Deep learning
Convolutional neural network
Fully convolutional neural network
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Mechanical Systems and Signal Processing 封面图
Mechanical Systems and Signal Processing
IF:
8.9
论文数:
1.3W
被引数:
6.6W

机构

K
korea research institute of standards & science (kriss)
学者数:
2.0K
论文数: 2.3K
被引数: 1
引用论文

引用论文

Rapid communication: sequencing of the porcine agouti-related protein (AGRP) gene
err2002-05-01
err0
PREAI
errM. J. Halverson; K. J. Donelan; N. H. Granholm; T. M. Cheesbrough; C. A. Westby; D. M. Marshall
err分享
err收藏
err分享
err收藏
学者 查看更多内容