arrow
Return

Dynamic Shields: A Game-Theoretic Reinforcement Learning Framework for APT Mitigation

delete2026-01-01
delete0
PRE
AI
G
Gustaf Johansson *
A
Aws Naser Jaber
F
Florian Skopik
M
Max Landauer
W
Wolfgang Hotwagner
M
Markus Wurzenberger
DOI:10.1007/978-3-032-08064-6_7delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Advanced Persistent Threats (APTs), exemplified by the SolarWinds attack, demand adaptive defenses in partially observable network environments. We model the defender-attacker interaction as a Partially Observable Markov Decision Process (POMDP)-based stochastic game, executed in the test environment AttackBed, and solved using reinforcement learning (RL) with Proximal Policy Optimization (PPO) and Recurrent PPO (RPPO). Our contributions include: (1) theorems proving equilibrium existence, threshold-structured best responses, and convergence properties, (2) a high-fidelity GNS3-based simulation aligned with MITRE ATT&CK/D3FEND frameworks, and (3) empirical comparisons showing PPO outperforms RPPO in mitigating attacks. PPO reduces attack success rates by 65%, leveraging sample efficiency in realistic settings. This comprehensive study advances game-theoretic RL for cyber defense, providing a foundation for future multi-agent frameworks.
Keywords:
Reinforcement Learning
Game Theory
Cybersecurity
Advanced PersistentThreats
POMDP
ThresholdPolicy

Journal

G
GAME THEORY AND AI FOR SECURITY, GAMESEC 2025, PT I
IF:
0
Papers:
16
Citations:
0

Organization

R
royal institute of technology
Scholars:
1.1K
Papers: 549
Citations: 0