Return
Exploring KV Cache Quantization in Multimodal Large Language Model Inference
DOI:10.1109/LCA.2025.3646170.png)
Abstract
En 中文
Multimodal large language models (MLLMs) have demon-strated strong performance across modalities, such as image, video, andaudio understanding, by leveraging large language models (LLMs) as abackbone. However, a critical challenge in MLLM inference is the largememory capacity required for the key-value (KV) cache, particularly whenprocessing high-resolutionimages. Thispressure oftenforcesheterogeneousCPU-GPU systems to offload the KV cache to CPU memory, introducingsubstantial transfer latency. KV cache quantization is a promising wayto reduce this memory demand, yet it remains underexplored for MLLMinference. In this work, we characterize MLLM inference and present atext-centric KV cache quantization method that retains only 10% of tokensin high precision while quantizing the rest. Our method reduces Time-To-First-Token (TTFT) by1.7xand Time-Per-Output-Token (TPOT) by4.3x, with negligible accuracy loss.
Keywords:
Quantization (signal)
Accuracy
Image resolution
Large language models
Graphics processing units
Sensitivity
Electric breakdown
Videos
Image sequences
Image reconstruction
Generative AI
multimodal large language models
quantization
Journal
I
IF:
1.4
Papers:
42
Citations:
781

