Return
Optimizing photographic composition with deep reinforcement learning
DOI:10.1016/j.neucom.2025.130363.png)
Abstract
En 中文
Image composition has always been a primary consideration in professional photography, with good composition significantly enhancing the visual appeal of an image. However, achieving rapid composition can be a challenge for most casual users. To address this issue, this paper proposes a view adjustment method to improve image composition. We introduce a DRL (Deep Reinforcement Learning) based approach for adjusting and optimizing image composition. This method employs D3QN (Dueling Double Deep Q-Network) as the main network architecture, leveraging the agent's progressive exploration to find the optimal policy for composition adjustments. Additionally, we have developed a composition feature extraction module that extracts compositional information from images according to the professional photography composition guidelines, thereby further enhancing the network's predictive performance. We also created a composition defect dataset for model training and testing, expanding on existing image cropping datasets. This dataset includes samples of displacement adjustment, scaling adjustment, and their combinations, aiding in the model's training. Through various experiments, our model demonstrated excellent performance, achieving an IoU (Intersection over Union) score of 0.785 and a Disp (Displacement) score of 0.050. The experimental results indicate that our proposed view adjustment method effectively improves image composition and can guide users in real-time to take wellcomposed photographs. In summary, this research aims to assist casual users in more easily obtaining professional-level composition guidance, utilizing intelligent technical means so that everyone can capture visually appealing moments.
Keywords:
Image composition
Composition defect dataset
Deep Reinforcement Learning
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

