Return
Text-Guided ROI-Based Neural Image Compression
DOI:10.1109/TBC.2025.3609065.png)
Abstract
En 中文
ROI-based image compression can significantly reduce data redundancy and improve storage efficiency while preserving the high fidelity of critical content. Current methods typically employ explicit bit allocation for region-adaptive coding by masking features before quantization to suppress background information. Although this strategy enhances compression performance, it severely degrades background reconstruction quality. In response, we propose a text-guided ROI-based neural image compression method, which mainly adopts a novel implicit bit allocation strategy similar to attention mechanism and incorporates natural-language descriptions to enable flexible user interaction. We firstly leverage a high-performance referencing-image segmentation model to generate masks from text prompts. Next, we introduce an efficient mask-guided feature transform (MGFT) module, composed of a region-adaptive attention (RAA) block and a region-adaptive transform (RAT) block, to more effectively apply these masks for implicit bit allocation. To improve reconstructed background quality, we design a weight-shared dual-decoder network that separately reconstructs foreground and background regions. To the best of our knowledge, this is the first work to integrate textual descriptions into ROI-based image compression and to employ implicit bit allocation for high-quality region-adaptive coding. Experiments on the COCO2017 dataset demonstrate that our method achieves optimal rate–distortion performance and supports text-specified targeted compression with excellent coding efficiency. The decoded results maintain high foreground fidelity while preserving pleasing background perceptual quality.
Keywords:
ROI-based image compression
mask-guided feature transform module
implicit bit allocation
region-adaptive attention block
region-adaptive transform block
Journal
IF:
4.8
Papers:
2.1K
Citations:
3.0K

