arrow
Return

CVML-Pose: Convolutional VAE Based Multi-Level Network for Object 3D Pose Estimation

delete2023-01-01
delete4
delete
OA
AI
Z
Zhao, Jianyu *
E
Edward Sanderson
B
Bogdan J. Matuszewski
DOI:10.1109/ACCESS.2023.3243551delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Most vision-based 3D pose estimation approaches typically rely on knowledge of object's 3D model, depth measurements, and often require time-consuming iterative refinement to improve accuracy. However, these can be seen as limiting factors for broader real-life applications. The main motivation for this paper is to address these limitations. To solve this, a novel Convolutional Variational Auto-Encoder based Multi-Level Network for object 3D pose estimation (CVML-Pose) method is proposed. Unlike most other methods, the proposed CVML-Pose implicitly learns an object's 3D pose from only RGB images encoded in its latent space without knowing the object's 3D model, depth information, or performing a post-refinement. CVML-Pose consists of two main modules: (i) CVML-AE representing convolutional variational autoencoder, whose role is to extract features from RGB images, (ii) Multi-Layer Perceptron and K-Nearest Neighbor regressors mapping the latent variables to object 3D pose including, respectively, rotation and translation. The proposed CVML-Pose has been evaluated on the LineMod and LineMod-Occlusion benchmark datasets. It has been shown to outperform other methods based on latent representations and achieves comparable results to the state-of-the-art, but without use of a 3D model or depth measurements. Utilizing the t-Distributed Stochastic Neighbor Embedding algorithm, the CVML-Pose latent space is shown to successfully represent objects' category and topology. This opens up a prospect of integrated estimation of pose and other attributes (possibly also including surface finish or shape variations), which, with real-time processing due to the absence of iterative refinement, can facilitate various robotic applications. Code available: https://github.com/JZhao12/CVML-Pose.
Keywords:
3D pose estimation
deep learning
variational autoencoder
synthetic data

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

U
University of Central Lancashire
Scholars:
2.9K
Papers: 2.9K
Citations: 3.5K
Cited Papers

Cited Papers

PoseRBPF: A Rao-Blackwellized Particle Filter for 6-D Object Pose Tracking
err2021-10-01
err84
errOAAI
errDeng, Xinke; Mousavian, Arsalan; Xiang, Yu; Xia, Fei; Bretl, Timothy; Fox, Dieter
errShare
errSave
EUV emission spectra in collisions of multiply charged Sn ions with He and Xe
err2010-03-03
err0
PREAI
errH Ohashi; S Suda; H Tanuma; S Fujioka; H Nishimura; A Sasaki; K Nishihara
errShare
errSave
Antiferromagnetic ruthenium(III)
err2002-05-01
err0
PREAI
errRichard L. Carlin; Ramon Burriel; Kenneth R. Seddon; Russell I. Crisp
errShare
errSave
errShare
errSave
Augmented Autoencoders: Implicit 3D Orientation Learning for 6D Object Detection
err2019-10-23
err85
PREAI
errSundermeyer, Martin; Marton, Zoltan-Csaba; Durner, Maximilian; Triebel, Rudolph
errShare
errSave
3D Object Recognition and Pose Estimation From Point Cloud Using Stably Observed Point Pair Feature
err2020-01-01
err22
errOAAI
errLi, Deping; Wang, Hanyun; Liu, Ning; Wang, Xiaoming; Xu, Jin
errShare
errSave
Public goods games in populations with fluctuating size
err2018-05-01
err0
errOAAI
errAlex McAvoy; Nicolas Fraiman; Christoph Hauert; John Wakeley; Martin A. Nowak
errShare
errSave
researcher View more