arrow
返回

ACTUAL: Audio Captioning With Caption Feature Space Regularization

delete2023-01-01
delete6
PRE
AI
Z
Zhang, Yiming
H
Hong Yu
R
Ruoyi Du
Z
Zheng‐Hua Tan
W
Wenwu Wang
马
马占宇 (Zhanyu Ma) *
Y
Yuan Dong
DOI:10.1109/TASLP.2023.3293015delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Audio captioning aims at describing the content of audio clips with human language. Due to the ambiguity of audio content, different people may perceive the same audio clip differently, resulting in caption disparities (i.e., the same audio clip may be described by several captions with diverse semantics). In the literature, the one-to-many strategy is often employed to train the audio captioning models, where a related caption is randomly selected as the optimization target for each audio clip at each training iteration. However, we observe that this can lead to significant variations during the optimization process and adversely affect the performance of the model. In this article, we address this issue by proposing an audio captioning method, named ACTUAL (Audio Captioning with capTion featUre spAce reguLarization). ACTUAL involves a two-stage training process: (i) in the first stage, we use contrastive learning to construct a proxy feature space where the similarities between captions at the audio level are explored, and (ii) in the second stage, the proxy feature space is utilized as additional supervision to improve the optimization of the model in a more stable direction. We conduct extensive experiments to demonstrate the effectiveness of the proposed ACTUAL method. The results show that proxy caption embedding can significantly improve the performance of the baseline model and the proposed ACTUAL method offers competitive performance on two datasets compared to state-of-the-art methods.
Keyword:
Audio captioning
contrastive learning
cross-modal task
caption consistency regularization

期刊

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
论文数:
2.6K
被引数:
1.1W

机构

L
Ludong University
学者数:
5.6K
论文数: 3.3K
被引数: 3.7K
B
beijing university of posts & telecommunications
学者数:
1.4W
论文数: 1.2W
被引数: 9
U
University of Surrey
学者数:
1.2W
论文数: 1.3W
被引数: 22
A
aalborg university
学者数:
1.6W
论文数: 1.7W
被引数: 22
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Local Information Assisted Attention-Free Decoder for Audio Captioning
err2022-01-01
err10
errOAAI
errXiao, Feiyang; Guan, Jian; Lan, Haiyan; Zhu, Qiaoxi; Wang, Wenwu
err分享
err收藏
Electrotransformation ofStreptococcus agalactiaewith plasmid DNA
err1994-06-01
err0
errOAAI
errM.Luisa Ricci; Riccardo Manganelli; Cesare Berneri; Graziella Orefici; Gianni Pozzi
err分享
err收藏
学者 查看更多内容