Return
Reinforcement Learning-Based Equipment Combination Selection Optimization for Multi-Layer Kill Webs
DOI:10.1109/tcc.2026.3683729.png)
Abstract
En 中文
Modern network-centric operations increasingly rely on multi-layer Kill Webs (KWs), enabling redundant and non-linear sensing-to-strike pathways while introducing a combinatorial equipment selection problem under uncertainty and resource constraints. This paper formulates a multi-layer KW equipment combination selection as a sequential decision-making problem by explicitly modeling heterogeneous equipment capabilities, resource constraints, and the network topology. To address this problem, we developed an RL learning-based optimization framework, where a multi-objective reward function integrates normalized relevance, operational risk, and timeliness, with a penalty mechanism for infeasible or incomplete kill-chain closure. Based on the jointly captured state information (e.g., network structure, equipment attributes, target characteristics, and resource availability), an Actor-Critic (AC) algorithm is developed to learn adaptive equipment combination selection across different operational stages using temporal-difference advantage estimation and entropy regularization. Simulation results under diverse battlefield scenarios demonstrate that the proposed framework consistently outperforms Deep Q-Network (DQN), Proximal Policy Optimization (PPO), and Particle Swarm Optimization (PSO), achieving at least a 19.6% improvement in overall operational effectiveness while maintaining low decision latency.
Keywords:
Kill web (KW)
kill-chain
multi-layer network
equipment combination selection
reinforcement learning
Journal
I
IF:
5
Papers:
1.8K
Citations:
4.3K

