Return
LiDAR-BIND-T: Temporally Consistent Sensor Modality Translation and Fusion for Robotic Applications
N
A
J
S
DOI:10.1109/tro.2026.3710400.png)
Abstract
En 中文
Robust autonomous navigation requires reliable perception when optical sensors fail under adverse conditions, such as fog, rain, or smoke. While radar and sonar provide complementary sensing capabilities, their inherent sparsity and noise present fundamental challenges for direct fusion with dense light detection and ranging (LiDAR) representations. Deep learning-based fusion of heterogeneous sensor modalities in a shared embedding space enables seamless translation from sparse measurements to dense LiDAR-like predictions. However, naive frame-independent fusion suffers from temporal inconsistencies, geometric flickering, and unstable features, which catastrophically degrade downstream simultaneous localization and mapping (SLAM) performance. To address this, we introduce a novel approach to temporal coherence through three complementary mechanisms: explicit temporal regularization encouraging smooth latent transitions, motion-aware transformation losses supervising perceptual displacement consistency, and learned temporal fusion that aggregates information across time. We formalize the latent temporal coherence condition as a desired condition and demonstrate systematic violations in baseline models that are eliminated by our method. Evaluation on real-world indoor datasets shows substantial SLAM improvements. We further contribute domain-adapted metrics (Fréschet video motion distance and correlation-peak distance) that correlate strongly with navigation performance. The framework maintains plug-and-play modularity while ensuring the temporal stability required for reliable autonomous systems.
Keywords:
Autonomous vehicle navigation
deep learning methods
robust/adaptive control
sensor fusion
SLAM
Journal
IF:
10.5
Papers:
3.3K
Citations:
2.8W
