Return
A Lightweight Frequency-Selection-Based Progressive Patch Transformer Network for Single Image Super-Resolution
DOI:10.1109/TCSVT.2026.3663242.png)
Abstract
En 中文
Recently, lightweight networks for single image super-resolution (SISR) have surged due to the need of resource-constrained devices, where divide-and-conquer multi-route model exhibits impressive trade-off between performance and computational cost. However, most existing divide-and-conquer multi-route models face two key limitations: 1) possible suboptimal decoupling of image components (e.g. smooth regions, edges and texture details) due to spatial-domain-only processing, and 2) inability to model global dependencies explicitly and capture structural information, hindering further performance gains. To address these drawbacks, we propose a lightweight frequency-selection-based progressive patch Transformer network (FSPPTN) for higher-quality SISR reconstruction. Specifically, we first propose a frequency selection module, in which we develop a frequency enhancement branch (FEB) to dynamically decouple different image components by introducing the window-based Fast Fourier transform (WFFT) and a learnable weight matrix, and a spatial restoration branch (SRB) to recalibrate and fuse cross-granularity features by designing a multi-gate mechanism for reconstructing the component information screened out by the FEB at current level. Secondly, we propose a lightweight multi-branch gradient-guided inter-patch self-attention to explicitly capture global structural similarities by summarizing structural information of each patch into a lower-dimensional space using the statistical properties of first-order gradients, thereby achieving explicit global dependencies modeling and lightweight. Extensive experimental results demonstrate that, in the vast majority of cases, FSPPTN outperforms state-of-the-art lightweight SISR methods in terms of both performance and computational overhead, especially for <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$\times 3$ </tex-math></inline-formula> and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$\times 4$ </tex-math></inline-formula> SR, e.g. FSPPTN outperforms MaIR-Small by even 0.14dB PSNR on Manga109 dataset for <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$\times 4$ </tex-math></inline-formula> SR even with 48.3% fewer parameters and 63.5% lower FLOPs. The code is available at: <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/yslyangshuli/FSPPTN-main</uri>
Keywords:
Single image super-resolution
lightweight
frequency selection
explicit global structural similarities
first-order gradients
Journal
IF:
11.1
Papers:
612
Citations:
3.1W

