arrow
Return

Infusing Multisource Heterogeneous Knowledge for Language-Conditioned Segmentation and Grasping

delete2024-01-01
delete0
PRE
AI
J
Jialong Xie
J
Jin Liu
Z
Zhenwei Zhu
王超群 cover
王超群 (Chaoqun Wang)
段苹 cover
段苹 (Peng Duan)
周风余 (Fengyu Zhou) *
DOI:10.1109/TIM.2024.3446625delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Language-conditioned segmentation and grasping (LCSG) requires the robot to simultaneously identify and grasp a specific object in accordance with human linguistic instruction. Existing methods generally involve semantic matching between the entire instruction and raw RGB images to ground the desired object, generating a grasp pose for that object. However, they overlook the semantic effectiveness of the nouns in the instruction, struggling with fine-grained object localization. Furthermore, they lack sufficient geometry context of the object to reason the optimal grasp pose, resulting in instability and failures in cluttered environments. In this article, we propose a knowledge-augmented refinement network (KARNet) to jointly conduct fine-grained object segmentation and grasping detection to tackle these challenges. Specifically, to mine the semantic context of the nouns, we introduce an entity semantic enhancement (ESE) module to fuse the knowledge from both the external knowledge base and the contrastive language-image pretraining (CLIP). Besides, a refinement decoder is proposed to generate segmentation masks and incorporate the geometry-aware features to yield suitable grasp poses for the desired objects. Notably, to achieve better context fusion between linguistics and vision, we further introduce a language-guided object parsing (LGOP) module to conduct coarse-to-fine multimodal fusion. We conduct extensive experiments on the cluttered household dataset and demonstrate that our proposed approach attains a grasping accuracy of 97.97% and segmentation OIoU of 94.35%, reaching state-of-the-art performance. The effectiveness of our method is further validated in real-world applications. The project and video can be found at https://karnetgrasp.github.io.
Keywords:
Grasping
Robots
Semantics
Visualization
Task analysis
Robot kinematics
Knowledge based systems
Grasp pose refinement
heterogeneous knowledge
language-conditioned grasping
pretrained models
referring image segmentation

Journal

IEEE Transactions on Instrumentation and Measurement cover
IEEE Transactions on Instrumentation and Measurement
IF:
5.9
Papers:
1.9W
Citations:
5.8W

Organization

L
Liaocheng University
Scholars:
7.8K
Papers: 6.1K
Citations: 8.8K
S
shandong university
Scholars:
9.3W
Papers: 6.4W
Citations: 94