Return
Evolutionary Convolutional Autoencoder for Remote Sensing Image Super-Resolution Reconstruction
Y
Z
L
Z
DOI:10.3788/LOP252149.png)
Abstract
En 中文
Objective High-resolution remote sensing imagery plays a crucial role in tasks such as land cover classification and object detection. However, existing sub-meter high-resolution imagery faces limitations in widespread application for continuous regional or global-scale observation due to high costs, low revisit rates, and meteorological or policy constraints. Consequently, medium-to-low-resolution data such as Landsat and Sentinel-2 remain dominant resources. These datasets often suffer from resolution inadequacies and quality degradation influenced by sensor performance and imaging conditions, thereby constraining their analytical potential. Super-resolution technology enables the reconstruction of high-resolution images without hardware upgrades, serving as a key method to enhance the practical value of medium-to-low-resolution remote sensing data. This is particularly critical in time-sensitive scenarios such as disaster response and dynamic monitoring. Although deep learning-based super-resolution methods have made significant progress, they still commonly suffer from issues like loss of high-frequency details, complex model structures, and difficulties in parameter tuning. Evolutionary deep learning, by integrating evolutionary algorithms with neural networks, offers a viable path for automating network architecture and hyperparameter search. Among these, genetic algorithms (GAs), with their robust global search capabilities and structural flexibility, are widely employed for automated network architecture design and optimization. They can explore efficient, task-specific network configurations with minimal human intervention. Currently, research directly applying GAs to super-resolution reconstruction of remote sensing images remains in its early stages. Considering that super-resolution networks are essentially a type of deep learning model for pixel-wise reconstruction, existing studies have confirmed the effectiveness of GAs in structural optimization for visual tasks such as image classification and face recognition. Therefore, they can be transferred to super-resolution tasks with adaptive modifications designed for reconstruction objectives. Specifically, integrating low-resolution remote sensing image input, detail recovery objectives, and training processes into a fitness function enables GAs to optimize critical structural elements and hyperparameters. These include network module connections, upsampling methods, channel counts, and loss function combinations. This approach overcomes the limitations of manual design, constructing super-resolution models that better align with the complex structures and multi-scale features inherent in remote sensing imagery. Methods To address common challenges in remote sensing image super-resolution tasks-such as insufficient feature extraction, loss of high-frequency details, and high computational overhead-this paper proposes a super-resolution reconstruction framework based on the adaptive evolutionary convolutional autoencoder (AECAE). Traditional fixed-structure models struggle to adapt their architecture to diverse terrain scenarios, limiting their ability to recover complex textural details. To overcome this problem, AECAE employs a GA to perform evolutionary search on the structure of the convolutional autoencoder (CAE). This enables the network to dynamically adjust based on feature characteristics, balancing multi-scenario adaptability, model lightweighting, and reconstruction quality. During evolution, the network structure and hyperparameters are encoded as chromosomes. Model architecture is iteratively optimized through fitness-driven selection, crossover, and mutation operations. This paper employs minimum validation loss as the fitness metric, training and evaluating candidate networks in each generation while retaining top-performing individuals for the next iteration. The crossover operation adopts a layer-index-based random exchange strategy to enhance structural diversity through reorganization. The mutation operation utilizes parameter perturbation and hierarchical addition/deletion mechanisms to effectively avoid local optima. As evolutionary iterations progress, the network structure gradually converges toward an optimal configuration, achieving efficient adaptation across diverse datasets and feature types. Addressing the significant scale variations and complex structural textures in remote sensing imagery, this paper further incorporates a multi-scale feature extraction (MSFE) module within the encoder. This module captures local structures and global semantic information across varying receptive fields through multi-branch, deeply-separable convolutions, significantly enhancing the model's ability to represent high-frequency textures. Its parallel convolutional layers employ distinct kernel sizes combined with ReLU activations to fuse cross-scale features, achieving a balance between detail preservation and artifact suppression. Results and Discussions To comprehensively evaluate the super-resolution performance of the proposed AECAE against existing advanced remote sensing image methods, we select models such as enhanced deep super-resolution network (EDSR), super-resolution using a generative adversarial network (SRGAN), super-resolution using very deep convolutional network (VDSR), super-resolution convolutional neural network (SRCNN), efficient sub-pixel convolutional neural network (ESPCN), and image restoration using Swin Transformer (SwinIR) for comparison. Simulation experiments are conducted on the AID dataset using peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) metrics; real-scene experiments are performed on the Songshan dataset using average gradient (AG) and natural image quality evaluator (NIQE) metrics. In remote sensing image processing, the proposed AECAE demonstrates significant advantages over other methods in texture detail restoration, effectively enhancing the detail quality of low-resolution images. For instance, comparing the visual effects of scenes 1 and 2, we observe that the images generated by AECAE appear more visually consistent with reality, exhibiting minimal visual detail loss. In contrast, comparative models like SRCNN and ESPCN struggle to recover severely degraded details. Furthermore, models like EDSR and SwinIR demonstrate markedly superior visual quality, however, despite higher overall fidelity, their results often appear overly smoothed, potentially leading to loss of fine textural details. Similarly, the visual comparison in scene 3 provides additional evidence of AECAE's advantages. These contrasts indicate that AECAE restores details more realistically than other methods, with buildings appearing clearer in the image. AECAE demonstrates robust super-resolution performance across most remote sensing scenarios, with average PSNR and SSIM values consistently exceeding those of other comparative models. Compared to alternatives, AECAE excels in handling image super-resolution tasks for real-world scenes. Close inspection reveals that while other methods exhibit noticeable blurring or oversmoothing, AECAE generates results with sharper texture details and edges. The image super-resolution process successfully restores authentic details in contours and achieves superior visual outcomes. These results validate AECAE's effectiveness and practicality in real-world applications. Quantitative comparisons reveal that AECAE achieves the lowest AG and NIQE values among all models, indicating its robust super-resolution performance in actual remote sensing scenarios. Conclusions In this paper, we design an AECAE-based super-resolution reconstruction method for remote sensing images, and introduce a MSFE module, named AECAE. Optimized using a GA, the model demonstrates superior performance on both simulated and real remote sensing datasets. We compare AECAE against six state-of-the-art methods (EDSR, SRGAN, VDSR, SRCNN, ESPCN, SwinIR) on both simulated and real datasets, conducting quantitative and qualitative evaluations on the real remote sensing dataset. AECAE achieves an average PSNR of 33.88 dB and an average SSIM of 0.9144, ranking first among all compared models. It also attains the optimal values for both NIQE and AG metrics. Additionally, AECAE demonstrates superior visual quality and minimal detail loss in texture restoration. This indicates that the proposed AECAE outperforms existing state-of-the-art models in super-resolution tasks, yielding results closer to real-world scenarios. However, AECAE exhibits certain limitations. The evolutionary optimization process incurs substantial computational costs. Future efforts should focus on optimizing the model to reduce computational expenses without compromising its effectiveness.
Keywords:
adaptive evolutionary convolutional autoencoder
multi-scale feature extraction
image super-resolution reconstruction
remote sensing
lightweight network
Journal
L
IF:
1
Papers:
505
Citations:
0

