arrow
Return

SAMAF: Sequence-to-sequence Autoencoder Model for Audio Fingerprinting

delete2020-05-22
delete5
PRE
AI
A
Abraham Báez-Suárez *
N
Nolan Shah
J
Juan A. Nolazco‐Flores
S
Shou‐Hsuan Stephen Huang
O
Omprakash Gnawali
W
Weidong Shi
DOI:10.1145/3380828delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Audio fingerprinting techniques were developed to index and retrieve audio samples by comparing a content-based compact signature of the audio instead of the entire audio sample, thereby reducing memory and computational expense. Different techniques have been applied to create audio fingerprints; however, with the introduction of deep learning, new data-driven unsupervised approaches are available. This article presents Sequence-to-Sequence Autoencoder Model for Audio Fingerprinting (SAMAF), which improved hash generation through a novel loss function composed of terms: Mean Square Error, minimizing the reconstruction error; Hash Loss, minimizing the distance between similar hashes and encouraging clustering; and Bitwise Entropy Loss, minimizing the variation inside the clusters. The performance of the model was assessed with a subset of VoxCelebl dataset, aspeech in the wild dataset. Furthermore, the model was compared against three baselines: Dejavu, a Shazam-like algorithm; Robust Audio Fingerprinting System (RAFS), a Bit Error Rate (BER) methodology robust to time-frequency distortions and coding/decoding transformations; and Panako, a constellation-based algorithm adding time-frequency distortion resilience. Extensive empirical evidence showed that our approach outperformed all the baselines in the audio identification task and other classification tasks related to the attributes of the audio signal with an economical hash size of either 128 or 256 bits for one second of audio.
Keywords:
Deep learning
sequence-to-sequence autoencoder
audio fingerprinting
audio identification
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

ACM Transactions on Multimedia Computing Communications and Applications cover
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
Papers:
2.0K
Citations:
5.4K

Organization

T
Tecnologico de Monterrey
Scholars:
7.6K
Papers: 5.7K
Citations: 5
U
university of houston system
Scholars:
1.4W
Papers: 1.4W
Citations: 16