Return
ESCAN: Efficient GPU sharing for cascade neural network inference
DOI:10.1016/j.neunet.2025.107703.png)
Abstract
En 中文
• We address the deployment challenges of cascaded models in GPU-sharing scenarios. • GPU sharing requires balancing parallel gains and redundant computation for cascade models. • We introduce ESCAN with batch-parallel execution and resource allocation. • ESCAN improves the efficiency of cascaded models by optimizing the execution mode and resource allocation. • ESCAN simplifies and enhances the efficiency of the optimization process.
Journal
IF:
6.3
Papers:
7.8K
Citations:
3.0W
Organization
No organization information available

