Return
Three-Dimensional Affordance Segmentation for Object Point Cloud Driven by Language Instructions
DOI:10.1631/ENG.ITEE.2026.0044.png)
Abstract
En 中文
The location where a robot grasps an object is closely related to the task type. For the same object, different user requirements may necessitate different grasping strategies. Visual affordance serves as a reliable source of prior knowledge for manipulation. Existing methods learn affordance from images or videos, but planar affordance lacks the spatial information required for 6-degree-of-freedom (6-DoF) manipulation. Furthermore, current approaches are limited to affordances associated with predefined categories and cannot directly infer affordances from user instructions. To address such limitations, we propose a novel task: instruction-driven three-dimensional (3D) object affordance segmentation. To support this research, we introduce an instruction-affordance dataset (IAD), a challenging dataset consisting of 7190 object instances across 20 common object categories, paired with 624 manipulation instructions that specify the corresponding affordances. To evaluate generalization to novel commands, our dataset includes both seen and unseen settings. Building on this, we design an instruction-driven 3D affordance segmentation (IDAS) network, which extracts point cloud features and integrates instruction features layer by layer. Given a user instruction, our method segments suggested manipulation regions on the object's point cloud, thereby guiding the selection of optimal grasp poses. Experimental results show that our method outperforms other related approaches under both seen and unseen settings, demonstrating generalization ability to diverse user commands and unknown affordances.
Keywords:
Payloads
MODIS
Feeds
Motion pictures
Contacts
Circuits and systems
Quadrature amplitude modulation
Videos
Modulation
Protocols
Visual affordance
Point cloud segmentation
Open vocabulary
Multimodal fusion
Service robot
Journal
E
IF:
0
Papers:
28
Citations:
0

