arrow
Return

Diffusion-Driven Self-Supervised Learning for Shape Reconstruction and Pose Estimation

delete2025-12-24
delete0
PRE
AI
J
Jingtao Sun
王耀南 cover
王耀南 (Yaonan Wang)
M
Mingtao Feng
C
Chao Ding
M
Mike Zheng Shou
A
Ajmal Mian
DOI:10.1109/TPAMI.2025.3647855delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Fully-supervised category-level pose estimation aims to determine the 6-DoF poses of unseen instances from known categories, requiring expensive manual labeling costs. Recently, various self-supervised category-level pose estimation methods have been proposed to reduce the requirement of the annotated datasets. However, most methods rely on synthetic data or 3D CAD model, and they are typically limited to addressing single-object pose problems without considering multi-objective tasks or shape reconstruction. To overcome these challenges and limitations, we introduce a diffusion-driven self-supervised network for multi-object shape reconstruction and categorical pose estimation, only leveraging the shape priors. Specifically, to capture the SE(3)-equivariant pose features and 3D scale-invariant shape information, we present a Prior-Aware Pyramid 3D Point Transformer. This module adopts a point convolutional layer with radial-kernels for pose-aware learning and a 3D scale-invariant graph convolution layer for object-level shape representation. Furthermore, we introduce a Pretrain-to-Refine Self-Supervised Training Paradigm to train our network. It enables proposed network to capture the associations between shape priors and observations, addressing the challenge of intra-class shape variations by utilising the diffusion mechanism. Extensive experiments conducted on four public datasets and a self-built dataset demonstrate that our method significantly outperforms state-of-the-art self-supervised category-level baselines and even surpasses some fully-supervised instance-level and category-level methods. The project page is released at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Self-SRPE</uri>.
Keywords:
Category-level pose estimation
shape reconstruction
3D transformer
diffusion model
self-supervised learning

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

T
the university of western australia
Scholars:
649
Papers: 323
Citations: 0
H
hunan university
Scholars:
4.4W
Papers: 3.3W
Citations: 70
N
National University of Singapore
Scholars:
7.5W
Papers: 6.4W
Citations: 11.4W
researcher View more organizations