Return
Self-Supervised Object Pose Estimation With Multitask Learning
DOI:10.1109/TCDS.2025.3571813.png)
Abstract
En 中文
Object pose estimation using learning-based methods often necessitates vast amounts of meticulously labeled training data. The process of capturing real-world object images under diverse conditions and annotating these images with 6 degrees of freedom (6DOF) object poses is both time-consuming and resource-intensive. In this study, we propose an innovative approach to monocular 6-D pose estimation through self-supervised learning, eliminating the need for labor-intensive manual annotations. Our method initiates by training a multitask neural network in a fully supervised manner, leveraging synthetic RGBD data. We leverage semantic segmentation, instance-level depth estimation, and vector-field prediction as auxiliary tasks to enhance the primary task of pose estimation. Subsequently, we harness advancements in multitask learning to further self-supervise the model using unlabeled real-world RGB data. A pivotal element of our self-supervised object pose estimation is a geometry-guided pseudolabel filtering module that relies on estimated depth from instance-level depth estimation. Our extensive experiments conducted on benchmark datasets demonstrate the effectiveness and potential of our approach in achieving accurate monocular 6-D pose estimation. Importantly, our method showcases a promising avenue for overcoming the challenges associated with the labor-intensive annotation process, offering a more efficient and scalable solution for real-world object pose estimation.
Keywords:
Multitask learning (MTL)
neural networks for development
object pose estimation (OPE)
visual system and development
Journal
IF:
4.9
Papers:
1.0K
Citations:
3.5K

