1
Return

Vision-based manipulation from single human video with open-world object graphs

delete2026-05-08
delete0
delete
OA
AI
Y
Yifeng Zhu *
A
Arisrei Lim *
P
Peter Stone
Y
Yuke Zhu
DOI:10.1007/s10514-026-10253-8delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
This work presents an object-centric approach to learning vision-based manipulation skills from human videos. We investigate the problem of robot manipulation via imitation in the open-world setting, where a robot learns to manipulate novel objects from a single video demonstration. We introduce ORION, an algorithm that tackles the problem by extracting an object-centric manipulation plan as an Open-World Object Graph from a single RGB or RGB-D video, and then deriving a policy that conditions on the resulting plan. Our method enables the robot to learn from videos captured by daily mobile devices and generalize to deployment environments with varying visual backgrounds, camera angles, spatial layouts, and novel object instances. We systematically evaluate our method on both short-horizon and long-horizon tasks, using RGB-D and RGB-only demonstration videos. Across our real-world evaluations on varied tasks and demonstration modalities (RGB-D / RGB), we observe an average success rate of 74.4%, demonstrating the efficacy of ORION in learning from a single human video in the open world. Additional materials can be found on the project website .
Keywords:
Robot Manipulation
Imitation From Human Videos
Object-Centric Learning
Open-World Imitation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Autonomous Robots cover
Autonomous Robots
IF:
4.3
Papers:
1.6K
Citations:
5.0K

Organization

U
University of Texas at Austin
Scholars:
1.1K
Papers: 510
Citations: 4
Cited Papers

Cited Papers

Citing Papers

Citing Papers