Return
SPGrasp: Spatiotemporal Prompt-Driven Grasp Synthesis in Dynamic Scenes
Y
H
W
Y
Z
G
陈
DOI:10.1109/JSTSP.2026.3671182.png)
Abstract
En 中文
Real-time interactive grasp synthesis for dynamic objects remains challenging, as existing instance-level methods struggle to achieve low-latency inference while maintaining robust temporal consistency. To bridge this gap, we propose SPGrasp (Spatiotemporal Prompt-driven dynamic Grasp synthesis), a novel framework that extends the Segment Anything Model 2 (SAM 2) for video-stream grasp estimation. Our core innovation integrates user prompts with a spatiotemporal context module, enabling real-time interaction with end-to-end latency as low as 59 ms while preserving consistent instance identity and grasp predictions in dynamic, cluttered scenes, including object overlap and occlusion. In benchmark evaluations, SPGrasp achieves instance-level grasp accuracies of 90.6% on OCID and 93.8% on Jacquard. On the GraspNet-1Billion dataset under continuous tracking, SPGrasp reaches 92.0% accuracy with 73.1 ms per-frame latency, corresponding to a 58.5% latency reduction over the prior state-of-the-art promptable method RoG-SAM while maintaining competitive accuracy. Real-world experiments further demonstrate reliable interactive grasping under frequent occlusions, achieving a 94.8% success rate. These results suggest that SPGrasp effectively mitigates the latency–interactivity trade-off in dynamic grasp synthesis.
Keywords:
Dynamic grasp synthesis
segment anything model
prompt-driven grasping
spatiotemporal context
Journal
IF:
13.7
Papers:
1.9K
Citations:
1.1W
