arrow
Return

Knowledge Graph-Driven Reinforcement Learning for Zero-Shot Vision-Language Navigation

delete2026-04-28
delete0
delete
OA
AI
Z
Zhang, Ye
Z
Zhao, Yandong *
L
Liu, He
T
Tengfei Shi
W
Weitao Jia
L
Li, Shenghong
DOI:10.3390/math14091485delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
To address the limitations of zero-shot generalization in Vision-Language Navigation (VLN), this paper proposes a novel knowledge graph-driven reinforcement learning approach. Our method constructs a hierarchical, dynamically updated knowledge graph online during the agent's real-time interaction with the environment, seamlessly aligning external semantic priors with continuous visual perception. By leveraging a Chain-of-Thought (CoT) prompting mechanism, the agent performs multi-hop reasoning to precisely locate target objects. Furthermore, we design an end-to-end optimized reinforcement learning framework that fuses multi-modal features and employs a task-oriented composite reward function. Extensive experiments in the AI2-THOR simulation environment demonstrate that the proposed method significantly improves navigation success rates in zero-shot settings. The results validate its robust generalization capabilities, particularly for unseen object categories and complex scene layouts.
Keywords:
embodied AI
Vision-Language Navigation
reinforcement learning
knowledge graph
Chain-of-Thought

Journal

Mathematics cover
Mathematics
IF:
2.2
Papers:
2.9K
Citations:
3.6W

Organization

T
taiyuan university of technology
Scholars:
6.5K
Papers: 2.0K
Citations: 0