arrow
Return

Self-Supervised Object Pose Estimation With Multitask Learning

delete2025-05-21
delete0
PRE
AI
D
Dinh-Cuong Hoang
P
Phan Xuan Tan
T
Tuan A. Duong
T
Tuan-Minh Huynh
D
Duc-Manh Nguyen
A
Anh-Nhat Nguyen
D
Duc-Long Pham
V
Van-Duc Vu
T
Thu-Uyen Nguyen
N
Ngoc-Anh Hoang
K
Khanh-Toan Phan
D
Duc-Thanh Tran
V
Van-Thiep Nguyen
N
Ngoc-Trung Ho
C
Cong-Trinh Tran
V
Van-Hiep Duong
DOI:10.1109/TCDS.2025.3571813delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Object pose estimation using learning-based methods often necessitates vast amounts of meticulously labeled training data. The process of capturing real-world object images under diverse conditions and annotating these images with 6 degrees of freedom (6DOF) object poses is both time-consuming and resource-intensive. In this study, we propose an innovative approach to monocular 6-D pose estimation through self-supervised learning, eliminating the need for labor-intensive manual annotations. Our method initiates by training a multitask neural network in a fully supervised manner, leveraging synthetic RGBD data. We leverage semantic segmentation, instance-level depth estimation, and vector-field prediction as auxiliary tasks to enhance the primary task of pose estimation. Subsequently, we harness advancements in multitask learning to further self-supervise the model using unlabeled real-world RGB data. A pivotal element of our self-supervised object pose estimation is a geometry-guided pseudolabel filtering module that relies on estimated depth from instance-level depth estimation. Our extensive experiments conducted on benchmark datasets demonstrate the effectiveness and potential of our approach in achieving accurate monocular 6-D pose estimation. Importantly, our method showcases a promising avenue for overcoming the challenges associated with the labor-intensive annotation process, offering a more efficient and scalable solution for real-world object pose estimation.
Keywords:
Multitask learning (MTL)
neural networks for development
object pose estimation (OPE)
visual system and development

Journal

IEEE Transactions on Cognitive and Developmental Systems cover
IEEE Transactions on Cognitive and Developmental Systems
IF:
4.9
Papers:
1.0K
Citations:
3.5K

Organization

F
FPT University
Scholars:
776
Papers: 439
Citations: 187
S
Shibaura Institute of Technology
Scholars:
1.5K
Papers: 1.3K
Citations: 969