arrow
Return

Error Resilient Online Reinforcement Learning Using Adaptive Statistical Checks

delete2025-01-01
delete0
PRE
AI
C
Chandramouli Amarnath
M
Mohamed Mejri
J
Jackson Isenberg
A
Abhijit Chatterjee
DOI:10.1109/TCAD.2025.3529820delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Online deep reinforcement learning (deep RL)-based systems are being increasingly deployed in a variety of safety-critical applications. Due to the dynamic nature of the environments they work in, onboard reinforcement learning (RL) hardware is vulnerable to soft errors from radiation, thermal effects and electrical noise that corrupts the results of computations. Existing approaches to on-line error resilience in machine learning systems have relied on the availability of large training datasets to configure resilience parameters. This is not always feasible for online RL systems. Similarly, other approaches involving specialized hardware or modifications to training algorithms are difficult to implement for onboard RL applications. In contrast, we present a novel error resilience approach for online RL that leverages running statistics of neuron output values collected across the (real-time) RL training process to configure error detection thresholds (called checks) for the deep RL forward pass. Similarly, we formulate checks on the deep RL backward pass using running statistical thresholds on reduced-dimension checksums of online learning weight updates to rapidly detect and correct errors in online deep RL training. In this methodology, statistical concentration bounds leveraging running statistics are used to diagnose neuron outputs or weights as erroneous. The use of running statistics allows the checks to adapt to changes caused by continual on-line RL training. Erroneous neurons are set to zero (suppressed) in the forward pass. Erroneous weight updates are frozen, allowing nonerroneous weight updates to proceed and allowing online learning without rerunning training episodes. Our approach is compared against the state of the art and validated on several RL algorithms as well as a hardware validation platform.
Keywords:
Deep neural networks (DNNs)
fault tolerance
machine learning
reinforcement learning (RL)
resilience
soft errors

Journal

I
IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
IF:
2.9
Papers:
586
Citations:
9.6K

Organization

G
Georgia Institute of Technology
Scholars:
1.8W
Papers: 1.4W
Citations: 5.9W