Return
Bayesian inference based learning automaton scheme in Q-model environments
DOI:10.1007/s10489-021-02230-8.png)
Abstract
En 中文
Learning automaton (LA) is a reinforcement learning unit that learns the optimal action in a stochastic environment. Great efforts have been made to improve the performance of LA in the environments that provide only reward or penalty. However, in many practical scenarios, the feedback from the environment splits into multiple levels. The later environment is recognized by the LA community as the Q-model. This paper studies the LA in Q-model environments, whose study has been scanty. We propose a novel Bayesian inference-based LA that is capable of functioning in Q-model environments, BILAML. We utilize Bayesian inference to estimate the environment's response to each action. Then, KL divergence metric is adopted for adaptive decision-making. The BILAML scheme is proved to be optimal and is evaluated to be superior to established LA frameworks by comprehensive experiments.
Keywords:
Learning automaton
Bayesian inference
Q-model environments
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
3.5
Papers:
7.5K
Citations:
1.7W

