arrow
Return

Process Adaptive Learning for Visual-Language Navigation

delete2026-01-01
delete0
PRE
AI
G
Gao, Chaoqi
B
Boyuan Zhang
Y
Yahong Han *
DOI:10.1007/978-3-032-04555-3_14delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Visual-language navigation (VLN) requires the agent to understand language instructions and navigate in unseen environments. The agent needs to align language instructions with real-time and historical features to achieve goal grounding. Existing methods struggle to align instructions with the environment and rely on rigid navigation processes, limiting performance in unseen environments. This stems from the disconnect between instructions and navigation processes, as well as the weak correlation between the endpoint and instructions, which interferes with navigation. Thus, we propose a Process-Adaptive cRoss-modal Transformer (PART), which dynamically adjusts the navigation history based on instructions and links real-time action prediction with memory reasoning. This approach enhances the alignment between instructions and environments while mitigating the adverse effects of overlapping trajectories. Additionally, PART uses a distance-adaptive loss function to reduce reliance on specific trajectories while reinforcing goal-directed learning, enhancing generalization in unseen environments. On the goal-oriented VLN benchmark REVERIE and the step-by-step VLN benchmark R2R, PART surpasses previous state-of-the-art methods.
Keywords:
Vision and language
Embodied AI
Agent
Navigation

Journal

A
ARTIFICIAL NEURAL NETWORKS AND MACHINE LEARNING-ICANN 2025, PT IV
IF:
0
Papers:
30
Citations:
0

Organization

T
tianjin university
Scholars:
7.9W
Papers: 5.7W
Citations: 88