返回
Flow-Based Reinforcement Learning
DOI:10.1109/ACCESS.2022.3209260.png)
摘要
En 中文
This paper presents a novel Flow-based reinforcement learning strategy to model agent systems that can adapt to complex and dynamic problem environments by incrementally mastering their skills. It is inspired by the psychological notion of Flow that describes the optimal mental state experienced by an individual when they are fully immersed in a task and find it intrinsically rewarding to engage with. The proposed model presents an algorithm to describe the Flow experience such that agents can be trained through finer distinctions to the challenges across training time to maintain them in the Flow zone. In contrast to the traditional and incremental learning approaches that suffer from limitations associated with overfitting, the Flow-based model drives agent behaviours not simply through external goals but also through intrinsic curiosity to improve their skills and thus the performance levels. Experimental evaluations are conducted across two simulation environments on a maze navigation task and a reward collection task with comparisons against a generic reinforcement learning model and an incremental reinforcement learning model. The results reveal that these two models are prone to overfit under different design decisions and loose the ability to perform in dynamic variations of the tasks in varying degrees. Conversely, the proposed Flow-based model is capable of achieving near optimal solutions with random environmental factors, appropriately utilising the previously learned knowledge to identify robust solutions to complex problems.
Keyword:
Machine learning
Adaptation models
Reinforcement learning
Psychology
Learning (artificial intelligence)
Complexity theory
Artificial intelligence
Flow
reinforcement learning
incremental learning
machine learning
artificial intelligence
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
暂无机构信息
引用论文
Zinc oxide films prepared by dc reactive magnetron sputtering at different substrate temperatures
Vacuum
IF0
Structure of the amino-terminal domain of Cbl complexed to its binding site on ZAP-70 kinase
Nature
IF0
Pseudo-rehearsal: Achieving deep reinforcement learning without catastrophic forgetting伪演练: 在没有灾难性遗忘的情况下实现深度强化学习
NEUROCOMPUTING
IF6.5

