返回
Graph-Based Audio Classification Using Pre-Trained Models and Graph Neural Networks
DOI:10.3390/s24072106.png)
摘要
En 中文
Sound classification plays a crucial role in enhancing the interpretation, analysis, and use of acoustic data, leading to a wide range of practical applications, of which environmental sound analysis is one of the most important. In this paper, we explore the representation of audio data as graphs in the context of sound classification. We propose a methodology that leverages pre-trained audio models to extract deep features from audio files, which are then employed as node information to build graphs. Subsequently, we train various graph neural networks (GNNs), specifically graph convolutional networks (GCNs), GraphSAGE, and graph attention networks (GATs), to solve multi-class audio classification problems. Our findings underscore the effectiveness of employing graphs to represent audio data. Moreover, they highlight the competitive performance of GNNs in sound classification endeavors, with the GAT model emerging as the top performer, achieving a mean accuracy of 83% in classifying environmental sounds and 91% in identifying the land cover of a site based on its audio recording. In conclusion, this study provides novel insights into the potential of graph representation learning techniques for analyzing audio data.
Keyword:
ecoacoustics
environmental sound classification
graph neural networks
graph representation learning
node classification
pre-trained models
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.5
论文数:
7.2W
被引数:
20.9W
机构
引用论文
Inter-symbol Interference in High Data Rate Transmit Reference UWB Transceivers高速数据率发射参考UWB收发器中的符号间干扰
Analysis and assessment of ship collision accidents using Fault Tree and Multiple Correspondence Analysis基于故障树和多重对应分析的船舶碰撞事故分析与评估

