arrow
返回

Distributed Policy Evaluation Under Multiple Behavior Strategies

delete2015-05-01
delete69
delete
OA
AI
S
Sergio Valcárcel Macua *
J
Jianshu Chen
S
Santiago Zazo
A
Ali H. Sayed
DOI:10.1109/TAC.2014.2368731delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
We apply diffusion strategies to develop a fully-distributed cooperative reinforcement learning algorithm in which agents in a network communicate only with their immediate neighbors to improve predictions about their environment. The algorithm can also be applied to off-policy learning, meaning that the agents can predict the response to a behavior different from the actual policies they are following. The proposed distributed strategy is efficient, with linear complexity in both computation time and memory footprint. We provide a mean-square-error performance analysis and establish convergence under constant stepsize updates, which endow the network with continuous learning capabilities. The results show a clear gain from cooperation: when the individual agents can estimate the solution, cooperation increases stability and reduces bias and variance of the prediction error; but, more importantly, the network is able to approach the optimal solution even when none of the individual agents can (e.g., when the individual behavior policies restrict each agent to sample a small portion of the state space).
Keyword:
Adaptive networks
Arrow-Hurwicz algorithm
diffusion strategies
distributed processing
gradient temporal difference
mean-square-error
reinforcement learning
saddle-point problem
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Automatic Control 封面图
IEEE Transactions on Automatic Control
IF:
7
论文数:
1.3W
被引数:
6.7W

机构

U
Universidad Politecnica de Madrid
学者数:
1.4W
论文数: 1.2W
被引数: 10
University of California System 封面图
University of California System
学者数:
37.5W
论文数: 33.7W
被引数: 6.6K
引用论文

引用论文

Combat-Related Traumatic Brain Injury and Its Implications to Military Healthcare
err2010-12-01
err0
PREAI
errKimberly S. Meyer; Donald W. Marion; Helen Coronel; Michael S. Jaffee
err分享
err收藏
Explicit and implicit reinforcement learning across the psychosis spectrum.
err2017-07-01
err0
errOAAI
errDeanna M. Barch; Cameron S. Carter; James M. Gold; Sheri L. Johnson; Ann M. Kring; Angus W. MacDonald; Diego A. Pizzagalli; J. Daniel Ragland; Steven M. Silverstein; Milton E. Strauss
err分享
err收藏
Resistance and functional training reduces knee extensor position fluctuations in functionally limited older adults
err2005-09-29
err0
PREAI
errTodd M. Manini; Brian C. Clark; Brian L. Tracy; Jeanmarie Burke; Lori Ploutz-Snyder
err分享
err收藏
err分享
err收藏
学者 查看更多内容