Return
nnWeb: Towards efficient WebGPU-based DNN inference via automatic collaborative offloading
DOI:10.1016/j.comnet.2025.111489.png)
Abstract
En 中文
In-browser neural network inference offers the promise of cross-platform AI applications, but faces severe latency and energy challenges on resource-constrained devices. In this paper, we present nnWeb, a WebGPU-based in-browser neural network inference framework with optimized latency and energy efficiency. nnWeb dynamically partitions neural network and facilitates the collaborative offloading between client browser and server. nnWeb operates in two phases: (1) layer-wise isolation-based profiling, which is used to predict per-layer execution latency and energy on heterogeneous hardware; and (2) asynchronous execution-based DNN partitioning, which continuously monitors network bandwidth and device load to select the optimal partition point using WebGPU’s native pipeline parallelism, minimizing total latency or energy consumption by solving a closed-form optimization at runtime. Extensive evaluation on various in-browser AI models and networking conditions shows that nnWeb achieves an average reduction of 30% to 52% in total inference latency compared with static partitioning. Moreover, nnWeb realizes energy savings ranging from 11.3% to 44.0% in contrast to standalone browser inference.
Journal
IF:
4.6
Papers:
1.5K
Citations:
1.6W
Organization
No organization information available

