Return
Accelerating Private Large Transformers Inference Through Fine-Grained Collaborative Computation
DOI:10.1109/TIFS.2025.3584639.png)
Abstract
En 中文
Homomorphic encryption (HE) and secret sharing (SS) enable computations on encrypted data, providing significant privacy benefits for large transformer-based models (TBM) in sensitive sectors like medicine and finance. However, private TBM inference incurs significant costs due to the coarse-grained application of HE and SS. We present <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">FASTLMPI</small>, a new approach to accelerate private TBM inference through fine-grained computation optimization. Specifically, through the fine-grained co-design of homomorphic encryption and secret sharing, <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">FASTLMPI</small> achieves efficient protocols for matrix multiplication, SoftMax, LayerNorm, and GeLU. In addition, <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">FASTLMPI</small> introduces a precise segmented approximation technique for differentiable non-linear functions, improving its fitting accuracy while maintaining a low polynomial degree. Compared to solution BOLT (S&P’24), <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">FASTLMPI</small> shows a remarkable 25.1% to 55.3% decrease in runtime and an impressive 39.0% reduction in communication costs.
Keywords:
Secure multiparty computation
homomorphic encryption
privacy preserving
large transformer models
Journal
IF:
8
Papers:
5.2K
Citations:
2.3W

