Return
Ra<italic>C</italic>: Robot Learning for Long-Horizon Tasks by Scaling <underline>R</underline>ecovery <underline>a</underline>nd <underline>C</underline>orrection
Z
R
N
J
R
Z
A
DOI:10.1109/tro.2026.3706552.png)
Abstract
En 中文
Modern paradigms for robotic learning train expressive policy architectures on large amounts of human demonstration data. Yet performance on contact-rich, deformable-object, and long-horizon tasks plateau far below perfect execution, even with thousands of expert demonstrations. This is due to the inefficiency of existing “expert” data collection procedures based on human teleoperation. To address this issue, we introduce RaC, a new phase of training on human-in-the-loop rollouts after imitation learning pretraining. In RaC, we fine-tune a robotic policy on human intervention trajectories that illustrate recovery and correction behaviors. Specifically, during a policy rollout, human operators intervene when failure appears imminent, first rewinding the robot back to a familiar, in-distribution state and then providing a corrective segment that completes the current subtask. Training on this data composition expands the robotic skill repertoire to include retry and adaptation behaviors, which we show are crucial for boosting both efficiency and robustness on long-horizon tasks. Across three real-world bimanual control tasks: shirt hanging, airtight container lid sealing, takeout box packing, and a simulated assembly task, RaC outperforms standard imitation learning by over 2×. To contextualize our results against the best known prior work, we note that RaC outperforms them while requiring roughly <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$\text{10}\times$</tex-math></inline-formula> less data collection time and fewer samples.
Keywords:
Human-in-the-loop
imitation learning
robotic manipulation
scaling analysis
Journal
IF:
10.5
Papers:
3.3K
Citations:
2.8W
