Return
Light Field Angular Super-Resolution Conditional VAE Model Constrained by Spatial-Angular Consistency
DOI:10.3788/AOS251094.png)
Abstract
En 中文
Objective In practical applications, traditional light field imaging faces multiple challenges such as the high cost of optical equipment, constraints on the physical dimensions of cameras, and a limited number of sensor pixels. These factors collectively make it particularly difficult to achieve high-resolution imaging simultaneously in both the angular and spatial dimensions. Existing light field angular super-resolution techniques overcome the limitations of traditional optical systems through various algorithmic approaches, enabling higher-quality imaging. However, current light field acquisition technologies exhibit an inherent trade-off between spatial and angular resolution, making it challenging to meet the requirements of high-precision applications. Leveraging spatial-angular consistency to achieve light field angular super-resolution constitutes a critical research problem in computational imaging. In this work, we utilize the spatial-angular coupling inherent in light field data to formulate spatial-angular consistency constrained mean squared error (MSE) loss and epipolar plane image (EPI) loss functions, aiming to optimize the geometric structure in the generated novel views. Methods To achieve light field angular super-resolution, we propose a spatial-angular consistency constrained conditional variational autoencoder (CVAE) model, which employs sparse viewpoint planes as input data and a disparity map as a conditional input. The core architecture of our network model comprises four components: an encoder, a decoder, a feature extraction module, and a disparity estimation module. Specifically, the encoder, constrained by the spatial-angular consistency of the 4D light field data, transforms the input data into the latent space, while the decoder, similarly constrained by this spatial-angular consistency, transforms the latent variables back into the original data space. Furthermore, to enforce the projection position constraints of the light field sub-aperture array, we design a viewpoint position-based MSE loss, and to address the linear geometric characteristics inherent in EPIs, we design an EPI loss. Moreover, we replace the Kullback- Leibler (KL) divergence traditionally used in VAEs with the maximum mean discrepancy (MMD) loss, which is employed in combination with the Wasserstein distance to measure the discrepancy between the latent variable distribution and the real data distribution. Results and Discussions To evaluate the performance of the proposed method, we compare it against two supervised deep learning approaches (DistgASR and LFEPICNN) and one unsupervised method (CVAE) on both synthetic and real-world light field datasets. Experiments are conducted for angular super-resolution from 3x3 to 7x7 and from 3x3 to 9x9 using the synthetic HCI-4D light field dataset, the Stanford real-world light field dataset, and a light field dataset captured by our research group. On the bedroom, bicycle, herbs, and origami scenes from the HCI-4D dataset, our method achieves higher PSNR and SSIM scores compared to the unsupervised CVAE algorithm, approaching the performance of supervised learning algorithms (Tables 3 and 5). To further verify the effectiveness and generalization capability of the proposed method, we compute the mean PSNR and SSIM over the 30 scenes and occlusion subsets of the Stanford dataset. Our algorithm achieves PSNR gains of 0.23 dB and 0.59 dB over the CVAE algorithm on these subsets, with corresponding SSIM gains (Table 7). Visual results for the cars scene (Fig. 9) and the occlusions16 scene (Fig. 10) from the Stanford dataset demonstrate that our method reconstructs structural details (e. g., car rearview mirrors) and the overall scene with fidelity comparable to supervised learning approaches. Furthermore, on the real-world bonsai light field data captured by our group, our method outperforms the supervised DistgASR method, achieving a PSNR gain of 1.28 dB and an SSIM improvement of 1.19% (Table 10). To test noise robustness, we introduce additive Gaussian noise at SNR levels of 30, 20, 10, and 5 dB into the Stanford Cars scene. The performance of the proposed method at SNR levels above 10 dB is nearly equivalent to that under noise-free conditions (Fig. 13). Ablation studies confirm the effectiveness of the individual modules: adding the disparity estimation module and the EPI loss, both independently and jointly, leads to consistent PSNR improvements (Table 13). A loss weight sensitivity analysis (Fig. 15) indicates that excessively low or high MMD weights lead to blurry reconstructions, while inappropriate EPI weights introduce geometric distortions, underscoring the importance of appropriately balanced loss weighting. Conclusions The proposed spatial-angular consistency-constrained conditional VAE model achieves high-quality angular super-resolution for sparse light field data. This is accomplished through the formulation of a sub-aperture projection position MSE loss and an EPI-based spatial-angular consistency loss. Experimental results indicate that the method achieves super-resolution performance comparable to supervised methods across various synthetic and real-world datasets, while also demonstrating strong generalization to unseen scenes. In future work, we will integrate spatial-angular consistency information from diverse parametric representations by combining the linear constraints of EPI parameterization, the internal consistency of macro-pixel parameterization, and the sub-aperture array parameterization into a more comprehensive model. By unifying these diverse constraints within a unified optimization framework, we expect to further improve the accuracy of light field angular super-resolution.
Keywords:
four-dimensional light field
spatial-angular consistency
angular super-resolution
conditional variational autoencoder model
Journal
A
IF:
1.6
Papers:
476
Citations:
4.7K
Organization
No organization information available

