arrow
Return

Efficient Learned Image Compression with Dual-space Aggregation Transformer

delete2025-12-02
delete0
PRE
AI
K
Kai Hu
井佩光 cover
井佩光 (Peiguang Jing)
R
Renhe Liu
B
Bo Wei
Y
Yu Liu
DOI:10.1016/j.dsp.2025.105797delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In recent years, learned image compression (LIC) techniques have gained significant success by leveraging Convolutional Neural Networks (CNNs) and Transformer architectures. Due to the complementary strengths of CNNs and Transformers in feature representation, these architectures have been widely integrated into LIC tasks using either cascaded or parallel collaborative structures. However, these collaborative approaches often lead to suboptimal rate-distortion performance or increased model complexity. To address this issue, this paper introduces a lightweight Dual-Space Aggregation Transformer (DSAT) module, which integrates a multi-scale convolutional block with an 11 × 11 large-scale convolution kernel into the Transformer block to adaptively aggregate local and global context features. Specifically, a gate perceptron (GP) block is employed to replace multi-layer perceptrons (MLPs), enabling adaptive focus on important features. Building upon the DSAT module, we design an efficient learned image compression method that achieves superior rate-distortion performance while maintaining lower model complexity. To enhance compression efficiency, we propose a Mixed Channel-Spatial Context (MCSC) entropy model, which accurately predicts latent symbols to reduce redundancy in both spatial and channel dimensions. Furthermore, to mitigate the impact of missing frequency information on coding performance, we introduce a novel non-local frequency loss into the LIC task and jointly optimize the method in both the frequency and pixel domains. Experimental results on prevalent datasets demonstrate that our approach delivers optimal rate-distortion performance. In particular, experiments on the Kodak dataset show a BD-rate gain of 9.52% over VVC, with lower complexity.

Journal

D
Digital Signal Processing
IF:
3
Papers:
653
Citations:
0

Organization

T
tianjin university
Scholars:
7.9W
Papers: 5.7W
Citations: 88
O
Okayama University
Scholars:
1.6W
Papers: 1.1W
Citations: 8.4K