arrow
Return

BEEMS: Boosting Machine Vision Efficiency via Computation Graph-Based Memory Smoothing

delete2026-01-01
delete0
PRE
AI
H
Hanjing Shen
F
Fangxin Liu *
L
Liu, Jian
L
Li Jiang
H
Haibing Guan
DOI:10.1145/3774934.3786430delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the rapid advances of deep learning-based computer vision (CV) technology, digital images are increasingly processed not by humans, but by downstream CV algorithms. In particular, the growing popularity of vision foundation models has heightened interest in deploying these models on edge devices. However, limited memory remains a key bottleneck, making memory footprint reduction essential. Mainstream model customization methods often require intensive deployment efforts and can severely degrade accuracy. Moreover, existing deep learning frameworks generally do not prioritize memory optimization. Existing memory management schemes face practical limitations, including layer-wise memory imbalance, high management overhead, and volatile memory budgets. To tackle these issues, this work focuses on compilationlevel optimizations that are explicitly designed to be memoryaware. We observe that memory usage during vision foundation model inference varies significantly over time (up to a 10x difference), with extended periods of low memory demand. Based on this, we propose BEEMS, a dual-objective compiler that optimizes both memory and latency by smoothing memory usage across the computational graph. BEEMS analyzes the vision foundation model computational graph to identify peak and trough operators in terms of memory demand. It then builds an efficient optimization search space, offering a flexible interface that applies different strategies based on operator characteristics. Specifically, peak operators are optimized using techniques such as operator partitioning, kernel substitution, swapping, and rematerialization to reduce memory pressure, while trough operators apply subgraph substitutions to improve latency. Experiments on six diverse models show that BEEMS reduces peak memory by up to 90% and improves latency by 10%, demonstrating its effectiveness in jointly optimizing memory and performance.
Keywords:
AI compilation
deep learning systems
memory optimization

Journal

P
PROCEEDINGS OF THE 31ST ACM SIGPLAN ANNUAL SYMPOSIUM ON PRINCIPLES AND PRACTICE OF PARALLEL PROGRAMMING, PPOPP 2026
IF:
0
Papers:
43
Citations:
0

Organization

S
shanghai jiao tong university
Scholars:
15.2W
Papers: 11.5W
Citations: 159
B
Beihang University
Scholars:
5.1W
Papers: 4.1W
Citations: 37
S
shanghai qi zhi institute
Scholars:
64
Papers: 45
Citations: 0
researcher View more organizations