arrow
Return

PPOAccel: A High-Throughput Acceleration Framework for Proximal Policy Optimization

delete2022-09-01
delete6
delete
OA
AI
Y
Yuan Meng *
S
Sanmukh R. Kuppannagari
R
Rajgopal Kannan
V
Viktor K. Prasanna
DOI:10.1109/TPDS.2021.3134709delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Reinforcement Learning (RL) is a major branch of AI that enables agents to learn optimal decision making via interaction with the environment. Proximal Policy Optimization (PPO) is the state-of-the-art policy optimization based RL algorithm which achieves superior overall performance on various benchmarks. A PPO agent iteratively optimizes its policy - a function which chooses optimal actions approximated by a DNN, with each iteration consisting of two computationally intensive phases: Sample Generation - where agents inference on its policy and interact with the environment to collect data, and Model Update - where the policy is trained using the collected data. In this paper, we develop the first high-throughput PPO accelerator on CPU-FPGA heterogeneous platform. Our unified systolic-array based design accelerates both the inference and the training of the deep neural network used in a RL algorithm, and is generalizable to various MLP and CNN models across a wide range of RL applications. We develop novel optimizations to simultaneously reduce data access and computation latencies, specifically: (a) optimal data flow mapping to systolic array, (b) novel memory-blocked data layout to enable streaming stall-free data access in both forward and backward propagations, and, (c) a systolic array compute sharing technique to mitigate load imbalance in the training of two networks. We evaluate our design on widely used robotics and gaming benchmarks, achieving 1.4x-26x and 1.3x-2.7x improvements in throughput, respectively, when compared with state-of-the-art CPU/CPU-GPU implementations.
Keywords:
Conferences
Portable document format
Indexes
Typesetting
Loading
Web sites
Warranties
Reinforcement learning
hardware accelerators
FPGA

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

U
university of southern california
Scholars:
4.6W
Papers: 3.8W
Citations: 51
United States Department of Defense cover
United States Department of Defense
Scholars:
2.8W
Papers: 2.3W
Citations: 172