arrow
返回

Model reusability in Reinforcement Learning

delete2025-05-12
delete0
delete
OA
AI
S
Sepideh Nikookar
S
Sohrab Namazi Nia
S
Senjuti Basu Roy *
S
Sihem Amer-Yahia
B
Behrooz Omidvar-Tehrani
DOI:10.1007/s00778-025-00920-0delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
在强化学习(Reinforcement Learning, RL)中,重复使用训练模型的能力对于复杂任务具有显著的实践价值。尽管在数据管理领域,监督模型的重复使用性已得到广泛研究,但据我们所知,这是首次为RL提出的基于原则的研究。为了捕捉训练策略,我们开发了一个基于表达性强且无损的图数据模型的框架,该框架能够容纳基于时序差分学习(Temporal Difference Learning)和深度强化学习(Deep-RL)的RL算法。我们的框架能够捕捉任意奖励函数,这些函数可以在推理时进行组合。该框架提供了理论保证,并表明其结果与从头训练的策略相同。我们设计了一个参数化算法,在累积奖励方面实现了效率与质量的平衡。在两个常见的RL任务(查询优化和机器人运动)上的实验验证了我们的理论,并展示了我们算法的有效性和效率。
Keyword:
Reinforcement Learning
Reusability of ML models
Optimization algorithms

期刊

VLDB Journal 封面图
VLDB Journal
IF:
3.8
论文数:
81
被引数:
2.4K

机构

U
university grenoble alpes
学者数:
315
论文数: 119
被引数: 1
N
New Jersey Institute of Technology
学者数:
4.2K
论文数: 4.5K
被引数: 4.6K
引用论文

引用论文

Cooperative Route Planning Framework for Multiple Distributed Assets in Maritime Applications多分布式资产在海上应用中的协同航路规划框架
err2022-06-11
err0
PREAI
errSepideh Nikookar; Paras Sakharkar; Sathyanarayanan Somasunder; Senjuti Basu Roy; Adam Bienkowski; Matthew Macesker; Krishna R. Pattipati; David Sidoti
err分享
err收藏
err分享
err收藏
err分享
err收藏
Informing sequential clinical decision-making through reinforcement learning: an empirical study
err2010-12-22
err136
errOAAI
errShortreed, Susan M.; Laber, Eric; Lizotte, Daniel J.; Stroup, T. Scott; Pineau, Joelle; Murphy, Susan A.
err分享
err收藏
err分享
err收藏
err分享
err收藏
Deep Reinforcement Learning: A Brief Survey深度强化学习: 简要综述
err2017-11-01
err0
errOAAI
errKai Arulkumaran; Marc Peter Deisenroth; Miles Brundage; Anil Anthony Bharath
err分享
err收藏
学者 查看更多内容