arrow
Return

<monospace>KaleidoScope</monospace>: A Co-Processor for Neural-Network-Driven Intelligent Data Plane

delete2026-04-29
delete0
PRE
AI
D
Dong Wen
Z
Zhongpei Liu
T
Tong Yang
李天匀 (Tianyun Li)
Y
Yanshu Wang
T
Tao Li
Z
Zhuochen Fan
Q
Qing Li
Z
Zhigang Sun
DOI:10.1109/tc.2026.3688717delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Neural-network-driven intelligent data plane (NN-driven IDP) is becoming an emerging topic. To be practically deployable in real-world networks, NN-driven IDP should satisfy three design goals: the generality to support various NNs models, the high performance with low latency and high throughput, and the transparency preserving the performance and functionality of data plane. However, it is challenging to simultaneously achieve both high performance and generality, and to accomplish low-overhead transparency. In this paper, we propose <monospace xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">KaleidoScope</monospace>, a data plane co-processor meets all design goals. Our key insight is that traffic analysis demands differential requirements: packet-level analysis needs low latency performance; while flow-level tasks require high accuracy. We therefore present a hybrid granularity inference engine, which leverages distinct architectures to support different NNs models, achieving both high performance and generality. To further ensure low-overhead transparency and cooperate with hybrid granularity inference, we design a lightweight traffic monitor, feature extraction, and inference scheduling engine. We prototype <monospace xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">KaleidoScope</monospace> on both FPGA and ASIC platforms. On FPGA, <monospace xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Kaleidoscope</monospace> achieves an inference latency as low as 256 ns and a throughput of 100 Gbps, with negligible impact on the data plane. On ASIC, it achieves over 1.1 Tbps throughput. The on-board tested NNs demonstrate state-of-the-art accuracy compared to existing work, highlighting the impact of generality on achieving high accuracy. Our code is open-source and available on GitHub: <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/Kaleidoscope-arch/Kaleidoscope</uri>.
Keywords:
Traffic analysis
neural network accelerator
data plane co-processor

Journal

IEEE Transactions on Computers cover
IEEE Transactions on Computers
IF:
3.8
Papers:
5.3K
Citations:
9.8K

Organization

N
national university of defense technology
Scholars:
4.3K
Papers: 1.4K
Citations: 0
P
pengcheng laboratory
Scholars:
440
Papers: 218
Citations: 0
P
peking university
Scholars:
11.7W
Papers: 8.7W
Citations: 146
researcher View more organizations