Return
SFTNet: Spatial-Frequency Transformer Network for Learned Image Compression
刘
L
J
姚
J
T
赵
DOI:10.1109/TBC.2025.3635389.png)
Abstract
En 中文
In learned image compression (LIC), most existing methods process spatial domain information through convolutional neural networks (CNNs) or Transformers. However, they struggle to model frequency-domain correlations, particularly high-frequency detail correlations across regions. They suffer from high-frequency detail redundancy and inefficient uniform bit allocation, neither of which adapts to regional complexity. To address these issues, we propose the Spatial-Frequency Transformer Network (SFTNet) for LIC, which models cross-domain correlations between spatial structures and frequency components to enable region-adaptive bit allocation. To model these correlations, the Frequency-aware Transformer Block (FATB) employs dual attention mechanisms to collaboratively process spatial and frequency components. Specifically, its intra-block frequency attention enhances high-frequency details within each block by using low-frequency components as local structural guidance. Meanwhile, its inter-block spatial attention captures global consistency by modeling cross-block dependencies among low-frequency components. Building on this, the Feature Re-weighting Strategy (FRS) evaluates block importance via joint analysis of spatial dependencies and high-frequency energy. This enables dynamic alignment of the spatial-frequency features modeled by FATB for adaptive bit allocation, prioritizing structurally complex regions and compressing smooth areas to reduce redundancy. Experimental results demonstrate that our SFTNet outperforms VVC by −6.55% and −6.18% in BD-Rate on Kodak and CLIC datasets while achieving better reconstruction quality of the high-frequency details compared to recent LIC methods.
Keywords:
Learned image compression
adaptive bit allocation
spatial-frequency features
dual attention mechanisms
Journal
IF:
4.8
Papers:
2.1K
Citations:
3.0K
