arrow
Return

Reconstructing Speech From CNN Embeddings

delete2021-01-01
delete7
PRE
AI
L
Luca Comanducci *
P
Paolo Bestagini
M
Marco Tagliasacchi
A
Augusto Sarti
S
Stefano Tubaro
DOI:10.1109/LSP.2021.3073628delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The complete understanding of the decision-making process of Convolutional Neural Networks (CNNs) is far from being fully reached. Many researchers proposed techniques to interpret what a network actually learns from data. Nevertheless many questions still remain unanswered. In this work we study one aspect of this problem by reconstructing speech from the intermediate embeddings computed by a CNNs. Specifically, we consider a pre-trained network that acts as a feature extractor from speech audio. We investigate the possibility of inverting these features, reconstructing the input signals in a black-box scenario, and quantitatively measure the reconstruction quality by measuring the word-error-rate of an off-the-shelf ASR model. Experiments performed using two different CNN architectures trained for six different classification tasks, show that it is possible to reconstruct time-domain speech signals that preserve the semantic content, whenever the embeddings are extracted before the fully connected layers.
Keywords:
Decoding
Task analysis
Spectrogram
Feature extraction
Computer architecture
Image reconstruction
Training
Audio processing
explainable deep learning
speech recognition
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Signal Processing Magazine cover
IEEE Signal Processing Magazine
IF:
9.6
Papers:
1.1W
Citations:
1.7W

Organization

A
alphabet inc.
Scholars:
1.1K
Papers: 663
Citations: 0
P
Polytechnic University of Milan
Scholars:
2.0W
Papers: 1.8W
Citations: 24