Return
Galaxy Redshift Estimation Using SDSS Multi-Band Photometric Images Based on Variational Autoencoder (Invited)
W
C
L
C
DOI:10.3788/LOP252273.png)
Abstract
En 中文
Objective Photometric redshift estimation is a crucial technique for modern large-scale cosmological surveys, providing distance measurements for vast numbers of galaxies. While traditional methods-such as template fitting and machine learning on photometric features-have been widely used, they often struggle with robustness and interpretability. Template fitting's accuracy is highly dependent on the representativeness of spectral energy distribution templates, while machine learning models typically rely on handcrafted features such as magnitudes and colors, potentially underutilizing the rich morphological information present in galaxy images. With the advent of deep learning, end-to-end approaches that directly process images have emerged. However, purely discriminative models often lack physically meaningful latent representations and may not generalize well. This study aims to develop a novel joint learning framework that integrates the representational power of generative models with the predictive accuracy of discriminative models. The primary objective is to achieve high-precision, robust, and interpretable photometric redshift estimation directly from multi-band galaxy images, effectively controlling outliers and providing insights into the learned features. Methods We propose a shared-encoder variational autoencoder (VAE) and regression joint learning framework. The model architecture consists of three main components: an encoder, a decoder, and a regression module. The encoder, shared by both the VAE and regression tasks, maps the input five-band (u, g, r, i, z) galaxy images from the Sloan digital sky survey (SDSS) data release 16 to parameters (mean and variance) of a latent Gaussian distribution. The latent variable is sampled using the reparameterization trick. The decoder attempts to reconstruct the original input image from this latent variable, while the regression module-a multi-layer perceptron-predicts the redshift value from the same latent variable. This dual-branch design forces the shared encoder to learn a latent representation that is both informative for reconstruction and predictive for redshift, thereby incorporating physical constraints. The model is trained using a composite loss function L-total = alpha L-reconstruction+ beta L-KL + L-regression. The reconstruction loss (L-reconstruction, mean squared error) ensures image fidelity. The Kullback-Leibler divergence loss (L-KL) regularizes the latent space to approximate a standard normal distribution. The regression loss (L-regression,Huber loss) focuses on robust redshift prediction. The weights (alpha, beta) are optimized, and a redshift-aware weighting strategy is employed during training to mitigate data imbalance. The data processing pipeline involves precise image registration, background noise estimation, source masking, conversion to the AB magnitude system, and min-max normalization, resulting in 128 pixel & times; 128 pixel, five-channel images. Results and Discussions The proposed method demonstrates superior performance on the test set. It achieves a normalized median absolute deviation (NMAD) of 0.0259, an abnormality rate of 0.6 %, and a normalized bias of 0.0023. These results outperform traditional template fitting methods and several machine learning baselines, particularly in controlling outliers. The plot of predicted versus true redshift shows a tight correlation along the one-to-one line, and the residual plot confirms the absence of significant systematic bias across most of the redshift range. Ablation studies validate the necessity of key components. Removing the redshift-aware loss increases the abnormality rate and introduces bias, while removing redshift weighting worsens the NMAD significantly, highlighting their roles in enhancing robustness and accuracy. The model's generative capability is evidenced by high-quality reconstructed images that retain key morphological features while effectively removing noise. Furthermore, the learned latent space exhibits meaningful structure. Clustering analysis using k-means on the principal component analysis (PCA)-reduced latent representation reveals three distinct clusters corresponding to low-, medium-, and high-redshift galaxies, indicating that the model autonomously organizes galaxies by distance. Crucially, a correlation analysis between latent variables and physical parameters from the MPA-JHU catalog reveals statistically significant Pearson correlations. Specific latent dimensions show strong correlations with stellar mass (e.g., Latent_56: r = 0.581), star formation rate (e.g., Latent_ 43: r =-0.262), and metallicity (e.g., Latent_29: r = - 0.311). This provides concrete evidence for the physical interpretability of the latent space, suggesting the model encodes fundamental galaxy properties. Performance on a high-redshift subsample (redshift over 0.4) is also analyzed. While metrics degrade (NMAD of 0.0291, abnormality rate of 2.8 % ) compared to the full test set, this is consistent with the challenge of data imbalance, as high-redshift galaxies are severely underrepresented in the training data. This analysis pinpoints the primary source of error for these objects. Conclusions This study successfully constructs and evaluates a shared-encoder VAE-regression joint learning framework for photometric redshift estimation. The experimental results lead to several key conclusions. First, the joint learning paradigm is effective, yielding redshift predictions that are more accurate and robust, with fewer outliers than several established methods. Second, the framework provides enhanced interpretability, as the latent space is demonstrated to encode information not only about redshift but also about intrinsic physical properties of galaxies, such as stellar mass and star formation rate. Finally, the careful design of the loss function, including task weighting and redshift-aware balancing, is crucial for stabilizing training and achieving high performance. The main limitation is the degraded performance on high-redshift galaxies, which stems directly from their scarcity in the training data. Future work will focus on employing active learning strategies to strategically augment the training set with high-redshift and rare galaxy types and exploring advanced multi-modal fusion techniques to further leverage the information in multi-band images.
Keywords:
photometric redshift
variational autoencoder
deep learning
galaxy image
Sloan digital sky survey
Journal
L
IF:
1
Papers:
505
Citations:
0

