Return
TensorMap: A Deep RL-Based Tensor Mapping Framework for Spatial Accelerators
DOI:10.1109/TC.2024.3398424.png)
Abstract
En 中文
The mapping of tensor computation is a complex and important process for spatial accelerators. Today's mapping works depend on hand-tuned kernel libraries or search-based heuristics from human experts. The former is time-intensive while the latter easily leads to sub-optimal performance. In this paper, we propose TensorMap, a deep reinforcement learning (RL)-based mapping framework for tensor computations on spatial accelerators. We propose a sequential generation mode for mapping optimization and construct a coarse-grained action space to reduce the complexity of the mapping search space. An efficient policy network is devised to optimize mapping primitives in the RL-based search. We then propose a stop signal that is sampled from Bernoulli distribution to facilitate multi-level loop unrolling for spatial accelerators. Finally, a genetic algorithm is employed to further refine the optimized mappings. In the experiments, we demonstrate TensorMap's ability for different spatial accelerators with various tensor computations. On TPU, TensorMap provides 2.6x, 2.7x, and 2.4x better energy-delay product (EDP) on average compared with FlexTensor, Ansor, and AMOS respectively. On Eyeriss, TensorMap provides 2.1 x, 1.8 x, and 1.7x better EDP on average compared with FlexTensor, Ansor, and AMOS respectively.
Keywords:
Tensors
Convolution
Computers
Parallel processing
Optimization
Libraries
Hardware
Tensor computation
spatial accelerator
software mapping
reinforcement learning
Journal
IF:
3.8
Papers:
5.3K
Citations:
9.8K

