arrow
Return

COAL: Robust Contrastive Learning-Based Visual Navigation Framework

delete2025-01-07
delete0
PRE
AI
Z
Zengmao Wang
胡建华 (Jianhua Hu)
Q
Q.R. Tang
高伟 cover
高伟 (Wei Gao) *
DOI:10.1002/rob.22508delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Real-world robots will face a wide variety of complex environments when performing navigation or exploration tasks, especially in situations where the robots have never been seen before. Usually, robots need to establish local or global maps and then use path planning algorithms to determine their routes. However, in some environments, such as a wild grassy path or pavement on either side of a road, it is difficult for robots to plan routes through navigation maps. To address this, we propose a robust framework for robot navigation using contrastive learning called Contrastive Observation–Action in Latent (COAL) space. To extract features from the action space and observation space, respectively, COAL uses two different encoders. At the training stage, COAL does not require any data annotation and a mask approach is employed to keep features with significant differences away from each other in latent space. Similar to multimodal contrastive learning, we maximize bidirectional mutual information to align the features of observations and action sequences in latent space, which can enhance the generalization of the model. At the deployment stage, robots only need the current image as observation to complete exploration tasks. The most suitable action sequence is selected from the sampled data for generating control signals. We evaluate the robustness of COAL in both simulation and real environments. Only 41 min of unlabeled training data is required to allow COAL to explore environments that have never been seen before, even at night. Compared with state-of-the-art methods, COAL has the strongest robustness and generalization ability. More importantly, the robustness of COAL is further improved by augmenting our training data using other open-source data sets, which indicates that our framework has great potential to extract deep features of observations and action sequences. Our code and trained models are available at https://github.com/wzm206/COAL.
Keywords:
contrastive learning
multimodal
self-supervised learning
visual navigation

Journal

Journal of Field Robotics cover
Journal of Field Robotics
IF:
5.2
Papers:
1.7K
Citations:
6.0K

Organization

I
Institute of Automation
Scholars:
539
Papers: 287
Citations: 220
U
University of Chinese Academy of Sciences
Scholars:
6.8K
Papers: 2.7K
Citations: 24.6W
Cited Papers

Cited Papers

University of Michigan North Campus long-term vision and lidar dataset
err2015-12-24
err300
PREAI
errCarlevaris-Bianco, Nicholas; Ushani, Arash K.; Eustice, Ryan M.
errShare
errSave
LaND: Learning to Navigate From Disengagements
err2021-04-01
err0
errOAAI
errGregory Kahn; Pieter Abbeel; Sergey Levine
errShare
errSave
Active Autonomous Aerial Exploration for Ground Robot Path Planning
err2017-04-01
err0
errOAAI
errJeffrey Delmerico; Elias Mueggler; Julia Nitsch; Davide Scaramuzza
errShare
errSave
GNM: A General Navigation Model to Drive Any Robot
err2023-05-29
err0
errOAAI
errDhruv Shah; Ajay Sridhar; Arjun Bhorkar; Noriaki Hirose; Sergey Levine
errShare
errSave
BotanicGarden: A High-Quality Dataset for Robot Navigation in Unstructured Natural Environments
err2024-03-01
err0
errOAAI
errYuanzhi Liu; Yujia Fu; Minghui Qin; Yufeng Xu; Baoxin Xu; Fengdong Chen; Bart Goossens; Poly Z.H. Sun; Hongwei Yu; Chun Liu; Long Chen; Wei Tao; Hui Zhao
errShare
errSave
MobileNetV2: Inverted Residuals and Linear Bottlenecks
err2018-06-01
err0
errOAAI
errMark Sandler; Andrew Howard; Menglong Zhu; Andrey Zhmoginov; Liang-Chieh Chen
errShare
errSave
Information based adaptive robotic exploration
err2024-09-19
err0
PREAI
errF. Bourgault; A.A. Makarenko; S.B. Williams; B. Grocholsky; H.F. Durrant-Whyte
errShare
errSave
researcher View more