Return
Toward Learned Image Compression for Multiple Semantic Analysis Tasks
DOI:10.1109/TBC.2025.3597154.png)
Abstract
En 中文
Deep neural network (DNN)-based image compression methods have demonstrated superior rate-distortion performance compared to traditional codecs in recent years. However, most existing DNN-based compression methods only optimize signal fidelity at certain bitrate for human perception, neglecting to preserve the richness of semantics in compressed bitstream. This limitation renders the images compressed by existing deep codecs unsuitable for machine vision applications. To bridge the gap between image compression and multiple semantic analysis tasks, an integration of self-supervised learning (SSL) with deep image compression is proposed in this work to learn generic compressed representations, allowing multiple computer vision tasks to perform semantic analysis from the compressed domain. Specifically, the semantic-guided SSL under bitrate constraint is designed to preserve the semantics of generic visual features and remove the redundancy irrelevant to semantic analysis. Meanwhile, a compression network with high-order spatial interactions is proposed to capture long-range dependencies with low complexity to remove global redundancy. Without incurring decoding cost of pixel-level reconstruction, the features compressed by the proposed method can serve multiple semantic analysis tasks in a compact manner. The experimental results from multiple semantic analysis tasks confirm that the proposed method significantly outperforms traditional codecs and recent deep image compression methods in terms of various analysis performances at similar bitrates. The source code of this work can be found in https://mic.tongji.edu.cn.
Keywords:
Image compression
self-supervised learning
machine vision
semantic analysis
Journal
IF:
4.8
Papers:
2.1K
Citations:
3.0K

