Return
ZSPose: Instance-Level Zero-Shot Object Pose Estimation With Segment Anything Model
DOI:10.1109/TASE.2025.3618671.png)
Abstract
En 中文
Estimating the poses of new objects is a challenging problem. Although many methods have been developed for instance-level object pose estimation, they often struggle when faced with new/unfamiliar objects. In this paper, we propose a zero-shot pose estimation method for new objects called ZSPose. We leverage SAM’s zero-shot feature to segment objects in cluttered environments and acquire masks for each object. To facilitate the matching of object masks with object models and categories, we propose a novel object matching strategy that aligns masks with the corresponding object models. Subsequently, based on the derived object masks, we produce object point clouds. Utilizing the RGB images of the objects alongside the point clouds, we present a feature-weight-based method for object pose estimation, achieving accurate pose estimation by predicting the matching weights between the model features and the point cloud features. We conduct performance testing on various instance-level object pose estimation datasets, and experimental results show that our proposed method significantly enhances the accuracy of object pose estimation. It demonstrates excellent generalization, making it applicable to pose estimation for a wide range of new objects. Finally, to validate the practical applicability of ZSPose, we apply it to real-world object pose estimation tasks and robotic grasping tasks. The experimental findings indicate that ZSPose effectively estimates the poses of new objects, assisting robots in performing practical grasping tasks, thus holding considerable practical value. Note to Practitioners—Instance-level object pose estimation is a crucial approach for accurately determining the poses of objects. It enables pose estimation tasks in various scenes when the information about the object model is available and performs well even in cluttered and occluded environments. Robots rely on this pose estimation to effectively complete object grasping tasks. However, most instance-level object pose estimation methods can only estimate the poses of objects that were encountered during the training phase, rendering them incapable of incorporating new objects. This limitation poses significant challenges for practical applications. In real industrial settings, the objects that robots need to grasp can vary widely and often include new items. The lack of scalability in current instance-level object pose estimation methods is a critical issue. To tackle this problem, we propose a novel zero-shot object pose estimation method called ZSPose. This method effectively estimates the poses of new objects that were not part of the training dataset and can be successfully applied to robotic grasping tasks.
Keywords:
Zero-shot
instance-level object pose estimation
segment anything model
Journal
IF:
6.4
Papers:
4.9K
Citations:
1.6W

