arrow
Return

ESCAN: Efficient GPU sharing for cascade neural network inference

delete2025-06-15
delete0
PRE
AI
W
Wang, Jianan
Y
Yang Shi
Z
Zhaoyun Chen
M
Mei Wen
DOI:10.1016/j.neunet.2025.107703delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We address the deployment challenges of cascaded models in GPU-sharing scenarios. • GPU sharing requires balancing parallel gains and redundant computation for cascade models. • We introduce ESCAN with batch-parallel execution and resource allocation. • ESCAN improves the efficiency of cascaded models by optimizing the execution mode and resource allocation. • ESCAN simplifies and enhances the efficiency of the optimization process.

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.8K
Citations:
3.0W

Organization

No organization information available