1
Return

Adaptive curriculum reinforcement learning with sim-to-real strategy in balance control of underactuated triple pendulum robots

delete2026-03-01
delete0
PRE
AI
Y
Yunfan Fu
J
Jing Guo
L
Li, Donghao
J
Junpeng Chen
L
Liangliang Han
Q
Qu, Yuanju
P
Pan, Yang *
DOI:10.1017/S0263574726103282delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper addresses the challenge of balance control for the underactuated triple pendulum robot (UTPR) using a model-free reinforcement learning (RL) strategy. A curriculum-based Soft Actor-Critic strategy, with a quadratic form and an integral term in the reward function (CSAC-QI), is proposed. By incorporating the integral of cumulative joint angle errors into the reward function, the CSAC-QI method significantly reduces steady-state errors and enhances control precision. CSAC-QI improves convergence efficiency through an adaptive curriculum learning (CL) framework that enables a structured transition from simpler to more complex tasks. To enhance control robustness, motor friction identification and domain randomization are implemented during training, thereby equipping the UTPR to cope with real-world uncertainties. Simulation experiments demonstrate superior performance of the CSAC-QI method in handling larger initial joint deviations, achieving accurate end-effector positioning, and maintaining balance under dynamic randomization, sensor noise, and external disturbances. Notably, the trained policy is directly deployed on the UTPR prototype, where it successfully maintains balance in real-world conditions.
Keywords:
underactuated robot
soft actor-critic
reward function
curriculum learning
balance control
sim-to-real

Journal

R
Robotica
IF:
2.7
Papers:
100
Citations:
4.1K

Organization

S
southern university of science & technology
Scholars:
1.0K
Papers: 394
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers