arrow
Return

Look-up-Table Based Processing-in-Memory Architecture With Programmable Precision-Scaling for Deep Learning Applications

delete2022-02-01
delete16
delete
OA
AI
P
Purab Ranjan Sutradhar
S
Sathwika Bavikadi
M
Mark Connolly
M
Mark Indovina
S
Sai Manoj Pudukotai Dinakarrao *
A
Amlan Ganguly
DOI:10.1109/TPDS.2021.3066909delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Processing in memory (PIM) architecture, with its ability to perform ultra-low-latency parallel processing, is regarded as a more suitable alternative to von Neumann computing architectures for implementing data-intensive applications such as Deep Neural Networks (DNN) and Convolutional Neural Networks (CNN). In this article, we present a Look-up Table (LUT) based PIM architecture aimed at CNN/DNN acceleration that replaces logic-based processing with pre-calculated results stored inside the LUTs in order to perform complex computations on the DRAM memory platform. Our LUT-based DRAM-PIM architecture offers superior performance at a significantly higher energy-efficiency compared to the more conventional bit-wise parallel PIM architectures, while at the same time avoids fabrication challenges associated with the in-memory implementation of logic circuits. Alongside, the processing elements can be programmed and re-programmed to perform virtually any operation, including operations of Convolutional, Fully Connected, Pooling, and Activating Layers of CNN/DNN. Furthermore, it is capable of operating on several combinations of bit-widths of the operand data and thereby offers a wider range of flexibility across performance, precision, and efficiency. Transmission Gate (TG) realization of the circuitry ensures minimal footprint from the PIM architecture. Our simulations demonstrate that the proposed architecture can perform AlexNet inference at a nearly 13x faster rate and 125x more efficiency compared to state-of-the-art GPU and also provides 1.35x higher throughput at 2.5x higher energy-efficiency than another recent DRAM-implemented LUT-based PIM architecture in its baseline operation mode. Moreover, it offers 12x higher frame-rate at 9x more efficiency per frame for the lowest operand precision setting, with respect to its own baseline operation mode.
Keywords:
Computer architecture
Random access memory
Table lookup
Performance evaluation
Registers
Parallel processing
Optimization
Processing in memory (PIM)
look-up table (LUT)
deep neural networks (DNN)
convolutional neural networks (CNN)

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

G
George Mason University
Scholars:
7.7K
Papers: 7.9K
Citations: 1.0W
R
Rochester Institute of Technology
Scholars:
3.7K
Papers: 3.3K
Citations: 45