Return
Error Resilient Online Reinforcement Learning Using Adaptive Statistical Checks
DOI:10.1109/TCAD.2025.3529820.png)
Abstract
En 中文
Online deep reinforcement learning (deep RL)-based systems are being increasingly deployed in a variety of safety-critical applications. Due to the dynamic nature of the environments they work in, onboard reinforcement learning (RL) hardware is vulnerable to soft errors from radiation, thermal effects and electrical noise that corrupts the results of computations. Existing approaches to on-line error resilience in machine learning systems have relied on the availability of large training datasets to configure resilience parameters. This is not always feasible for online RL systems. Similarly, other approaches involving specialized hardware or modifications to training algorithms are difficult to implement for onboard RL applications. In contrast, we present a novel error resilience approach for online RL that leverages running statistics of neuron output values collected across the (real-time) RL training process to configure error detection thresholds (called checks) for the deep RL forward pass. Similarly, we formulate checks on the deep RL backward pass using running statistical thresholds on reduced-dimension checksums of online learning weight updates to rapidly detect and correct errors in online deep RL training. In this methodology, statistical concentration bounds leveraging running statistics are used to diagnose neuron outputs or weights as erroneous. The use of running statistics allows the checks to adapt to changes caused by continual on-line RL training. Erroneous neurons are set to zero (suppressed) in the forward pass. Erroneous weight updates are frozen, allowing nonerroneous weight updates to proceed and allowing online learning without rerunning training episodes. Our approach is compared against the state of the art and validated on several RL algorithms as well as a hardware validation platform.
Keywords:
Deep neural networks (DNNs)
fault tolerance
machine learning
reinforcement learning (RL)
resilience
soft errors
Journal
I
IF:
2.9
Papers:
586
Citations:
9.6K

