返回
Increasing sample efficiency in deep reinforcement learning using generative environment modelling
DOI:10.1111/exsy.12537.png)
摘要
En 中文
Reinforcement learning is a broad scheme of learning algorithms that, in recent times, has shown astonishing performance in controlling agents in environments presented as Markov decision processes. There are several unsolved problems in current state-of-the-art that causes algorithms to learn suboptimal policies, or even diverge and collapse completely. Parts of the solution to address these issues may be related to short- and long-term planning, memory management and exploration for reinforcement learning algorithms. Games are frequently used to benchmark reinforcement learning algorithms as they provide a flexible, reproducible and easy to control environments. Regardless, few games feature the ability to perceive how the algorithm performs exploration, memorization and planning. This article presents The Dreaming Variational Autoencoder with Stochastic Weight Averaging and Generative Adversarial Networks (DVAE-SWAGAN), a neural network-based generative modelling architecture for exploration in environments with sparse feedback. We present deep maze, a novel and flexible maze game-engine that challenges DVAE-SWAGAN in partial and fully observable state-spaces, long-horizon tasks and deterministic and stochastic problems. We show results between different variants of the algorithm and encourage future study in reinforcement learning driven by generative exploration.
Keyword:
artificial experience-replay
deep reinforcement learning
environment Modelling
exploration
generative adversarial networks
generative Modelling
Markov decision processes
model-based RL
neural networks
variational autoencoder
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
2.3
论文数:
2.5K
被引数:
3.8K
机构
引用论文
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Fabrication of porous hollow γ-Al2O3 nanofibers by facile electrospinning and its application for water remediation静电纺丝法制备多孔中空 γ-Al2O3纳米纤维及其在水体修复中的应用
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning在mdp和半mdp之间: 强化学习中的时间抽象框架
A Deep Deterministic Policy Gradient-Based Method for Enforcing Service Fault-Tolerance in MEC基于深度确定性策略梯度强化服务容错的方法在移动边缘计算中的应用

