arrow
返回

Increasing sample efficiency in deep reinforcement learning using generative environment modelling

delete2020-03-01
delete4
delete
OA
AI
P
Per‐Arne Andersen *
M
Morten Goodwin
O
Ole‐Christoffer Granmo
DOI:10.1111/exsy.12537delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Reinforcement learning is a broad scheme of learning algorithms that, in recent times, has shown astonishing performance in controlling agents in environments presented as Markov decision processes. There are several unsolved problems in current state-of-the-art that causes algorithms to learn suboptimal policies, or even diverge and collapse completely. Parts of the solution to address these issues may be related to short- and long-term planning, memory management and exploration for reinforcement learning algorithms. Games are frequently used to benchmark reinforcement learning algorithms as they provide a flexible, reproducible and easy to control environments. Regardless, few games feature the ability to perceive how the algorithm performs exploration, memorization and planning. This article presents The Dreaming Variational Autoencoder with Stochastic Weight Averaging and Generative Adversarial Networks (DVAE-SWAGAN), a neural network-based generative modelling architecture for exploration in environments with sparse feedback. We present deep maze, a novel and flexible maze game-engine that challenges DVAE-SWAGAN in partial and fully observable state-spaces, long-horizon tasks and deterministic and stochastic problems. We show results between different variants of the algorithm and encourage future study in reinforcement learning driven by generative exploration.
Keyword:
artificial experience-replay
deep reinforcement learning
environment Modelling
exploration
generative adversarial networks
generative Modelling
Markov decision processes
model-based RL
neural networks
variational autoencoder
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Expert Systems 封面图
Expert Systems
IF:
2.3
论文数:
2.5K
被引数:
3.8K

机构

U
University of Agder
学者数:
2.1K
论文数: 2.3K
被引数: 3.4K
引用论文

引用论文

Deep Reinforcement Learning: A Brief Survey深度强化学习: 简要综述
err2017-11-01
err2.4K
errOAAI
errArulkumaran, Kai; Deisenroth, Marc Peter; Brundage, Miles; Bharath, Anil Anthony
err分享
err收藏
SOPHIE velocimetry of Kepler transit candidates XVII. The physical properties of giant exoplanets within 400 days of period
err2016-02-17
err187
errOAAI
errSanterne, A.; Moutou, C.; Tsantaki, M.; Bouchy, F.; Hebrard, G.; Adibekyan, V.; Almenara, J. -M.; Amard, L.; Barros, S. C. C.; Boisse, I.; Bonomo, A. S.; Bruno, G.; Courcol, B.; Deleuil, M.; Demangeon, O.; Diaz, R. F.; Guillot, T.; Havel, M.; Montagnier, G.; Rajpurohit, A. S.; Rey, J.; Santos, N. C.
err分享
err收藏
学者 查看更多内容