Return
Vision-based manipulation from single human video with open-world object graphs
Y
A
P
Y
DOI:10.1007/s10514-026-10253-8.png)
Abstract
En 中文
This work presents an object-centric approach to learning vision-based manipulation skills from human videos. We investigate the problem of robot manipulation via imitation in the open-world setting, where a robot learns to manipulate novel objects from a single video demonstration. We introduce ORION, an algorithm that tackles the problem by extracting an object-centric manipulation plan as an Open-World Object Graph from a single RGB or RGB-D video, and then deriving a policy that conditions on the resulting plan. Our method enables the robot to learn from videos captured by daily mobile devices and generalize to deployment environments with varying visual backgrounds, camera angles, spatial layouts, and novel object instances. We systematically evaluate our method on both short-horizon and long-horizon tasks, using RGB-D and RGB-only demonstration videos. Across our real-world evaluations on varied tasks and demonstration modalities (RGB-D / RGB), we observe an average success rate of 74.4%, demonstrating the efficacy of ORION in learning from a single human video in the open world. Additional materials can be found on the project website .
Keywords:
Robot Manipulation
Imitation From Human Videos
Object-Centric Learning
Open-World Imitation
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
4.3
Papers:
1.6K
Citations:
5.0K
