arrow
Return

MetaKernel: Enabling Efficient Encrypted Neural Network Inference through Unified MVM and Convolution

delete2025-10-01
delete0
PRE
AI
P
Peng Yuan *
Y
Yan Liu
J
J. W. Lai
L
Long Li
T
Tianxiang Sui
X
X. D. Zhang
Q
Qing Zhu
J
Jingling Xue
DOI:10.1145/3763095delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Practical encrypted neural network inference under the CKKS fully homomorphic encryption (FHE) scheme relies heavily on accelerating two key kernel operations: Matrix-Vector Multiplication (MVM) and Convolution (Conv). However, existing solutions-such as expert-tuned libraries and domain-specific languages-are designed in an ad hoc manner, leading to significant inefficiencies caused by excessive rotations. We introduce MKR, a novel composition-based compiler approach that optimizes MVM and Conv kernel operations for DNN models under CKKS within a unified framework. MKR decomposes each kernel into composable units, called MetaKernels, to enhance SIMD parallelism within ciphertexts (via horizontal batching) and computational parallelism across them (via vertical batching). Our approach tackles previously unaddressed challenges, including reducing rotation overhead through a rotation-aware cost model for data packing, while also ensuring high slot utilization, uniform handling of inputs with arbitrary sizes, and compatibility with the output tensor layout. Implemented in a production-quality FHE compiler, MKR achieves inference time speedups of 10.08x-185.60x for individual MVM and Conv kernels and 1.75x-11.84x for end-to-end inference compared to a state-of-the-art FHE compiler. Moreover, MKR enables homomorphic execution of large DNN models, where prior methods fail, significantly advancing the practicality of FHE compilers.
Keywords:
FHE
CKKS
FHE Compilers
MetaKernel

Journal

P
Proceedings of the ACM on Programming Languages-PACMPL
IF:
2.8
Papers:
308
Citations:
4.7K

Organization

A
ant group
Scholars:
235
Papers: 117
Citations: 0