arrow
返回

Efficient Incremental Offline Reinforcement Learning With Sparse Broad Critic Approximation

delete2024-01-01
delete6
PRE
AI
L
Liang Yao
B
Baoliang Zhao *
徐鑫 封面图
徐鑫 (Xin Xu)
Z
Ziwen Wang
P
Pak Kin Wong
胡莹 封面图
胡莹 (Ying Hu)
DOI:10.1109/TSMC.2023.3305498delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Offline reinforcement learning (ORL) has been getting increasing attention in robot learning, benefiting from its ability to avoid hazardous exploration and learn policies directly from precollected samples. Approximate policy iteration (API) is one of the most commonly investigated ORL approaches in robotics, due to its linear representation of policies, which makes it fairly transparent in both theoretical and engineering analysis. One open problem of API is how to design efficient and effective basis functions. The broad learning system (BLS) has been extensively studied in supervised and unsupervised learning in various applications. However, few investigations have been conducted on ORL. In this article, a novel incremental ORL approach with sparse broad critic approximation (BORL) is proposed with the advantages of BLS, which approximates the critic function in a linear manner with randomly projected sparse and compact features and dynamically expands its broad structure. The BORL is the first extension of API with BLS in the field of robotics and ORL. The approximation ability and convergence performance of BORL are also analyzed. Comprehensive simulation studies are then conducted on two benchmarks, and the results demonstrate that the proposed BORL can obtain comparable or better performance than conventional API methods without laborious hyperparameter fine-tuning work. To further demonstrate the effectiveness of BORL in practical robotic applications, a variable force tracking problem in robotic ultrasound scanning (RUSS) is investigated, and a learning-based adaptive impedance control (LAIC) algorithm is proposed based on BORL. The experimental results demonstrate the advantages of LAIC compared with conventional force tracking methods.
Keyword:
Broad learning system (BLS)
incremental critic design
linear critic approximation (LCA)
offline reinforcement learning (ORL)
variable force tracking

期刊

IEEE Transactions on Cybernetics 封面图
IEEE Transactions on Cybernetics
IF:
10.5
论文数:
1.1W
被引数:
5.0W

机构

S
shenzhen institute of advanced technology, cas
学者数:
5.6K
论文数: 4.5K
被引数: 7
U
University of Macau
学者数:
1.1W
论文数: 1.3W
被引数: 2.0W
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
学者 查看更多内容