返回
Toward deep drum source separation
DOI:10.1016/j.patrec.2024.04.026.png)
摘要
En 中文
In the past, the field of drum source separation faced significant challenges due to limited data availability, hindering the adoption of cutting -edge deep learning methods that have found success in other related audio applications. In this letter, we introduce StemGMD, a large-scale audio dataset of isolated single -instrument drum stems. Each audio clip is synthesized from MIDI recordings of expressive drum performances using ten real -sounding acoustic drum kits. Totaling 1224 h, StemGMD is the largest audio dataset of drums to date and the first to comprise isolated audio clips for every instrument in a canonical nine -piece drum kit. We leverage StemGMD to develop LarsNet, a novel deep drum source separation model. Through a bank of dedicated U -Nets, LarsNet can separate five stems from a stereo drum mixture faster than real-time and is shown to considerably outperform state-of-the-art nonnegative spectro-temporal factorization methods.
Keyword:
Deep learning
Drums
Music decomposition
Source separation
U-Net
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

