arrow
Return

Multimodal Pretrained Knowledge for Real-world Object Navigation

delete2025-06-26
delete0
PRE
AI
H
Hui Yuan
黄岩 (Yan Huang)
N
Naigong Yu *
张东波 (Dongbo Zhang)
Z
Zetao Du
Z
Ziqi Liu
K
Kun Zhang
DOI:10.1007/s11633-024-1537-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Most visual-language navigation (VLN) research focuses on simulate environments, but applying these methods to real-world scenarios is challenging because of misalignments between vision and language in complex environments, leading to path deviations. To address this, we propose a novel vision-and-language object navigation strategy that uses multimodal pretrained knowledge as a cross-modal bridge to link semantic concepts in both images and text. This improves navigation supervision at key-points and enhances robustness. Specifically, we 1) randomly generate key-points within a specific density range and optimize them on the basis of challenging locations; 2) use pretrained multimodal knowledge to efficiently retrieve target objects; 3) combine depth information with simultaneous localization and mapping (SLAM) map data to predict optimal positions and orientations for accurate navigation; and 4) implement the method on a physical robot, successfully conducting navigation tests. Our approach achieves a maximum success rate of 66.7%, outperforming existing VLN methods in real-world environments.
Keywords:
Visual-and-language object navigation
key-points
multimodal pretrained knowledge
optimal positions and orientations
physical robot

Journal

Machine Intelligence Research cover
Machine Intelligence Research
IF:
8.7
Papers:
301
Citations:
882

Organization

S
School of Automation and Electronic Information
Scholars:
23
Papers: 11
Citations: 0
S
School of Information Science and Technology
Scholars:
447
Papers: 158
Citations: 0
researcher View more organizations