arrow
Return

Robust sequential decision-making in adversarial environments

delete2026-05-23
delete0
delete
OA
AI
J
Jurij Ružejnikov *
T
Tatiana V. Guy
DOI:10.1080/21642583.2026.2646376delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Reinforcement learning (RL) agents often fail in adversarial environments where the Markov Decision Process (MDP) assumption of a stationary environment is violated. While model-free solutions for this setting exist, planning-based counterparts remain less explored. This paper introduces offline and online value iteration algorithms within the Threatened Markov Decision Process (TMDP) framework, in which the RL agent maintains and updates a Bayesian belief over the adversary's policy. The belief is integrated into a modified Bellman optimality equation to compute robust policies. We evaluate our framework with the stochastic adversarial multi-agent Coin Game. Our primary finding is that the model-based agent outperforms the TMDP version of model-free Q-learning by a significant margin, confirming that the benefits of model-based planning extend from MDP to TMDP. Furthermore, the proposed framework maintains a performance advantage over Q-learning baselines even when the system's transition function is unknown. The RL agent also demonstrated robustness to direct adversarial interactions. This work validates TMDP value iteration as an effective, planning-based approach for decision-making against adaptive adversaries.
Keywords:
Dynamic programming
adversarial machine learning
multi-agent reinforcement learning
robust reinforcement learning
Bayesian reinforcement learning

Journal

S
Systems Science & Control Engineering
IF:
0
Papers:
39
Citations:
0

Organization

C
Czech Academy of Sciences
Scholars:
2.7K
Papers: 1.1K
Citations: 4.4W
Cited Papers

Cited Papers

errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
Learning to cooperate against ensembles of diverse opponents
err2025-08-01
err0
PREAI
errPerera,Isuri; de Nijs,Frits; García,Julian
errShare
errSave
On a general concept of forgetting
err1993-10-01
err0
PREAI
errR. KULHAVÝ; M. B. ZARROP
errShare
errSave
researcher View more