返回
6-D Object Pose Estimation Using Multiscale Point Cloud Transformer
DOI:10.1109/TIM.2022.3222467.png)
摘要
En 中文
Predicting 6-D object pose is an essential task in vision measurement for robotic manipulation. RGB-based methods have a natural disadvantage due to the lack of 3-D information, thus leading to inferior results. Therefore, exploiting the geometry information in depth images is crucial to achieve accurate predictions. To this end, we propose a multiscale point cloud transformer (MSPCT) to better learn point cloud feature representations. MSPCT mainly consists of three types of modules: local transformer (LT), DownSampling (DS) module, and global transformer (GT). Specifically, the LT is designed to dynamically divide a local region centered on each point and further extract point-level feature with local context awareness. The DS module is utilized to decrease the resolution and enlarge the receptive field. GT is employed to capture global-range dependencies between extracted local features. Based on the proposed transformer blocks, we design a network architecture for object pose estimation, where we further obtain multiscale features by fusing the local features from LT and global features from GT to predict object's pose. Extensive experiments verify the effectiveness of LT and GT, and our pose estimation pipeline achieves promising results on three benchmark datasets.
Keyword:
3-D deep learning
object pose estimation
point cloud
self-attention
transformer
期刊
IF:
5.9
论文数:
2.0W
被引数:
5.8W
机构
引用论文
Inflexibility of mental planning: A characteristic disorder with prefrontal lobe lesions?心理计划的僵化: 前额叶病变的特征性障碍?
EUV emission spectra in collisions of multiply charged Sn ions with He and Xe多电荷Sn离子与He和Xe碰撞中的EUV发射光谱

