Return
Causal learning with uncertainty-aware transformer for vision-and-language navigation
DOI:10.1016/j.neucom.2025.132196.png)
Abstract
En 中文
Vision-and-Language Navigation (VLN), as an important research direction in embodied artificial intelligence, has attracted extensive attention in recent years due to its great potential in real-world applications. However, existing VLN methods inevitably learn spurious correlations and struggle to handle various uncertainties during navigation, which leads to biased predictions and poor generalization performance. To address this problem, in this paper, we propose a novel probabilistic model named Uncertainty-aware Causal Transformer (UCT), which is based on the theory of uncertainty-aware causal learning. This model is designed to train a powerful VLN agent that can eliminate confounders causing spurious correlations in uncertain environments, thereby learning unbiased feature representations. Specifically, following the back-door adjustment paradigm in causal learning, we design two debiasing modules for language and vision respectively: Text Uncertainty Causal Attention (TUCA) and Vision Uncertainty Causal Attention (VUCA). These modules eliminate confounders and establish genuine causal relationships based on the uncertainty modeling in each modality. To better balance model’s reliability and generalization, we further propose a training strategy based on uncertainty measurement, implemented with a High-Uncertainty Module (HUM) and a Low-Uncertainty Module (LUM), which takes into account the impact of the inherent uncertainty of data on the model during training. Experimental results show that our method significantly outperforms previous state-of-the-art approaches.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W
Organization
No organization information available

