arrow
Return

Computer Audition: From Task-Specific Machine Learning to Foundation Models

delete
delete0
delete
OA
AI
A
Andreas Triantafyllopoulos
I
Iosif Tsangko
A
Alexander Gebhard
A
Annamaria Mesaros
T
Tuomas Virtanen
B
Björn W. Schuller
DOI:10.1109/JPROC.2025.3593952delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Foundation models (FMs) are increasingly spearheading recent advances on a variety of tasks that fall under the purview of computer audition—i.e., the use of machines to understand sounds. They feature several advantages over traditional pipelines: among others, the ability to consolidate multiple tasks in a single model, the option to leverage knowledge from other modalities, and the readily available interaction with human users. Naturally, these promises have created substantial excitement in the audio community and have led to a wave of early attempts to build new, generalpurpose FMs for audio. In the present contribution, we give an overview of computational audio analysis as it transitions from traditional pipelines toward auditory FMs. Our work highlights the key operating principles that underpin those models and showcases how they can accommodate multiple tasks that the audio community previously tackled separately.
Keywords:
Acoustic scene classification
artificial intelligence (AI)
audio captioning (AC)
computational audio analysis
computer audition
foundation models (FMs)
large audio models
machine listening
sound event detection (SED)

Journal

Proceedings of the IEEE cover
Proceedings of the IEEE
IF:
25.9
Papers:
9.9K
Citations:
4.5W

Organization

T
tum university hospital
Scholars:
236
Papers: 77
Citations: 0
T
Tampere University
Scholars:
1.4W
Papers: 1.3W
Citations: 1.4W