arrow
Return

Three-Dimensional Affordance Segmentation for Object Point Cloud Driven by Language Instructions

delete2026-04-01
delete0
PRE
AI
D
Du, Jiaxuan
W
Wu, Hao *
T
Tian, Guohui
Z
Zhao, Zhixian
L
Leng, Shuwen
DOI:10.1631/ENG.ITEE.2026.0044delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The location where a robot grasps an object is closely related to the task type. For the same object, different user requirements may necessitate different grasping strategies. Visual affordance serves as a reliable source of prior knowledge for manipulation. Existing methods learn affordance from images or videos, but planar affordance lacks the spatial information required for 6-degree-of-freedom (6-DoF) manipulation. Furthermore, current approaches are limited to affordances associated with predefined categories and cannot directly infer affordances from user instructions. To address such limitations, we propose a novel task: instruction-driven three-dimensional (3D) object affordance segmentation. To support this research, we introduce an instruction-affordance dataset (IAD), a challenging dataset consisting of 7190 object instances across 20 common object categories, paired with 624 manipulation instructions that specify the corresponding affordances. To evaluate generalization to novel commands, our dataset includes both seen and unseen settings. Building on this, we design an instruction-driven 3D affordance segmentation (IDAS) network, which extracts point cloud features and integrates instruction features layer by layer. Given a user instruction, our method segments suggested manipulation regions on the object's point cloud, thereby guiding the selection of optimal grasp poses. Experimental results show that our method outperforms other related approaches under both seen and unseen settings, demonstrating generalization ability to diverse user commands and unknown affordances.
Keywords:
Payloads
MODIS
Feeds
Motion pictures
Contacts
Circuits and systems
Quadrature amplitude modulation
Videos
Modulation
Protocols
Visual affordance
Point cloud segmentation
Open vocabulary
Multimodal fusion
Service robot

Journal

E
ENGINEERING INFORMATION TECHNOLOGY & ELECTRONIC ENGINEERING
IF:
0
Papers:
28
Citations:
0

Organization

S
shandong university
Scholars:
9.3W
Papers: 6.4W
Citations: 94