Return
Future object localization using multi-modal ego-centric video
DOI:10.1016/j.jvcir.2025.104684.png)
Abstract
En 中文
Future object localization (FOL) seeks to predict the future locations of objects using information from past and present video frames. Ego-centric videos from vehicle-mounted cameras serve as a key source. However, these videos are constrained by a limited field of view and susceptibility to external conditions. To address these challenges, this paper presents a novel FOL approach that combines ego-centric video data with point cloud data, enhancing both robustness and accuracy. The proposed model is based on a deep neural network that prioritizes front-camera ego-centric videos, exploiting their rich visual cues. By integrating point cloud data, the system improves three-dimensional (3D) object localization. Furthermore, the paper introduces a novel method for ego-motion prediction. The ego-motion prediction network employs multi-modal sensors to comprehensively capture physical displacement in both 2D and 3D spaces, effectively handling occlusions and the limited perspective inherent in ego-centric videos. Experimental results indicate that the proposed FOL system with ego-motion prediction (MS-FOLe) outperforms existing methods on large-scale open datasets for intelligent driving.
Journal
IF:
3.1
Papers:
529
Citations:
5.6K

