arrow
Return

FD-DeCap: A Front-Door Causal Inference-Based Framework for Debiasing Automatic Audio Captioning

delete2026-01-01
delete0
PRE
AI
J
Jinyun Liu
H
Hui Li
M
Mingjun Wei
Z
Zhanlin Ji *
Z
Zhang, Haiyang *
И
Иван Ганчев *
DOI:10.1109/ACCESS.2026.3651636delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Automatic Audio Captioning (AAC) aims at generating natural language descriptions for audio content. However, existing methods are often affected by latent confounders and spurious co-occurrence patterns in the data, leading to bias and semantic inaccuracies. This paper proposes FD-DeCap, a front-door causal inference-based framework, for the AAC task. The framework consists of three core components: 1) an AudioAug module introduces noise perturbations in audio features to enhance robustness against environmental interference; 2) a MedGate module explicitly introduces a mediator variable to satisfy the identifiability conditions of the front-door criterion, thereby disentangling direct and indirect effects; and 3) a MSeCE consistency loss jointly optimizes cross-entropy and MSE constraints, encouraging reliance on mediator representations rather than spurious correlations. Experimental results demonstrate that FD-DeCap achieves stable performance improvements, compared to state-of-the-art frameworks, on the Clotho and AudioCaps datasets, with SPIDEr scores of 0.282 and 0.429, respectively. A multi-perspective causal validation of the front-door adjustment, performed on the Clotho dataset, includes analyzes of similarity-score distributions, feature distributions, and representative case studies. After debiasing, the similarity between generated captions and reference captions shifts upward overall, the mediator feature distributions become more dispersed, and the representative cases more accurately capture true acoustic scenes. These findings indicate that the proposed FD-DeCap framework effectively alleviates bias caused by latent confounders and spurious co-occurrence, enhances semantic consistency and robustness of generated captions, and provides a novel solution for the AAC task in complex acoustic scenarios.
Keywords:
Semantics
Birds
Feature extraction
Transformers
Training
Dogs
Data models
Convolutional neural networks
Chirp
Acoustics
Automatic audio captioning (AAC)
bias
causal inference
front-door adjustment

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

U
university of limerick
Scholars:
1.2K
Papers: 564
Citations: 0
Z
Zhejiang A&F University
Scholars:
1.0W
Papers: 6.1K
Citations: 178
N
north china university of science & technology
Scholars:
6.6K
Papers: 3.7K
Citations: 5
X
xi'an jiaotong-liverpool university
Scholars:
1.0K
Papers: 536
Citations: 0
researcher View more organizations