返回
Distributed Robust Bandits With Efficient Communication
DOI:10.1109/TNSE.2022.3231320.png)
摘要
En 中文
The Distributed Multi-Armed Bandit (DMAB) is a powerful framework for studying many network problems. The DMAB is typically studied in a paradigm, where signals activate each agent with a fixed probability, and the rewards revealed to agents are assumed to be generated from fixed and unknown distributions, i.e., stochastic rewards, or arbitrarily manipulated by an adversary, i.e., adversarial rewards. However, this paradigm fails to capture the dynamics and uncertainties of many real-world applications, where the signal that activates an agent, may not follow any distribution, and the rewards might be partially stochastic and partially adversarial. Motivated by this, we study the asynchronously stochastic DMAB problem with adversarial corruptions where the agent is activated arbitrarily, and rewards initially sampled from distributions might be corrupted by an adversary. The objectives are to simultaneously minimize the regret and communication cost, while robust to corruption. To address all these issues, we propose a Robust and Distributed Active Arm Elimination algorithm, namely RDAAE, which only needs to transmit one real number (e.g., an arm index, or a reward) per communication. We theoretically prove that the performance of regret and communication cost smoothly degrades when the corruption level increases.
Keyword:
Stochastic processes
Costs
Autonomous aerial vehicles
Adaptation models
Uncertainty
Indexes
Heuristic algorithms
Distributed multi-agent bandit (DMAB)
Adversarial corruptions
Cooperation
Robust learning
期刊
I
IF:
7.9
论文数:
2.6K
被引数:
10.0K
机构
引用论文
LCD Monitors as an Alternative for Precision Demanding Visual Psychophysical Experiments
Perception
IF0
Distributed Machine Learning for Wireless Communication Networks: Techniques, Architectures, and Applications无线通信网络的分布式机器学习: 技术,体系结构和应用
Communication-Efficient and Distributed Learning Over Wireless Networks: Principles and Applications
PROCEEDINGS OF THE IEEE
IF25.9
Decentralized Multi-Agent Multi-Armed Bandit Learning With Calibration for Multi-Cell Caching具有多单元缓存校准的分散式多智能体多臂Bandit学习

