Return
Lossless graph transformations for memory-efficient on-device machine learning
DOI:10.1016/j.sysarc.2026.104000.png)
Abstract
En 中文
Efficient memory management is a critical requirement for on-device machine learning, where limited memory capacity often becomes a key bottleneck. While existing model compression techniques successfully reduce model size, they intrinsically incur information loss and lead to accuracy degradation. In this work, we propose general lossless graph transformations that improve memory efficiency without sacrificing accuracy. We first identify and formulate three structural patterns of computational graphs that commonly contribute to excessive peak memory consumption. Based on the patterns, we design lossless graph optimizations that systematically eliminate them. We implement the proposed graph transformations on top of the ONNX-MLIR compiler framework and demonstrate up to 67.9% reduction in peak memory usage across representative on-device machine learning models.
Keywords:
On-device machine learning
Computational graph transformation
Memory optimization
Multi-level intermediate representation
Journal
IF:
4.1
Papers:
3.0K
Citations:
4.2K
Organization
No organization information available
Cited Papers
No cited papers available

