Return
Data driven computer vision oriented image signal processing pipeline configuration
DOI:10.1016/j.neucom.2026.134761.png)
Abstract
En 中文
The image signal processing (ISP) pipeline in camera systems is traditionally designed for human vision. A human vision-oriented ISP pipeline is optimized independently of the backend vision algorithms for computer vision systems, which may lead to a suboptimal solution for the entire system. To find the optimal configuration of the ISP pipeline for vision tasks, in this paper, we propose a hybrid supervised learning and reinforcement learning-based framework to co-design these two parts in a unified framework. To achieve this, an ISP policy network, ISPNet, is designed to output the vision performance-oriented ISP configuration with a given input RAW image for a specific computer vision task. In particular, ISPNet is trained together with a vision task model, where the output policy aims to maximize a vision performance-related reward while bypassing as many ISP steps as possible (for overall computational complexity reduction). Extensive experiments are conducted on the basis of classification and object detection. The results show that the entire ISP pipeline, or most ISP stages, can be eliminated if the following consumer is computer vision tasks instead of human eyes. Models trained on RAW images are robust enough for tasks such as classification and object detection under most normal image shooting conditions. It is also shown that the demosaicing step is necessary for tasks with dedicated pre-processing and/or data augmentation since it can resist interpolation artifacts and improve the generalization ability of models. While under degraded conditions (i.e., simulated low-light with additive Gaussian noise), the proposed dynamic configuration framework can outperform a fixed policy in terms of model performance.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

