Return
Probabilistic hypergraphs using multiple randomly masked autoencoders for semi-supervised multi-modal multi-task learning
DOI:10.1016/j.neucom.2026.134870.png)
Abstract
En 中文
The computer vision domain has greatly benefited from an abundance of data across many modalities or interpretation layers, in order to improve on various visual tasks. Recently, there has been a lot of focus on self-supervised pretraining methods through Masked Autoencoders (MAE) He et al. (2022) [1] and Bachmann et al. (2022) [2], usually used as a first step before optimizing for a given downstream task, such as classification or regression. The approach proves to be highly efficient and useful in learning as it does not require any manually labeled data. In this work, we introduce Probabilistic Hypergraphs using Masked Autoencoders (PHG-MAE): a novel model using multiple modalities or interpretation layers of the input data or tasks, which unifies recent work on multi-layer neural hypergraphs [3–6] with the widespread approach of MAEs under a common computational framework. Through random masking of entire modalities, not just patches of those layers, we show that the model samples from the distribution of hyperedges on each forward pass. Additionally, the model adapts the standard MAE algorithm by combining pretraining and fine-tuning into a single training loop. Moreover, our approach enables the creation of inference-time ensembles which, through aggregation, boost the final prediction performance and consistency. Lastly, we show that we can apply knowledge distillation on top of the ensembles with little loss in performance, even with models that have fewer than 1M parameters. While our experiments are focused on outdoor UAV scenes, the same type of tests can be performed in virtually any domain where multiple data layers and modalities are available. In order to streamline the process of integrating external pretrained experts for computer vision multi-modal multi-task learning (MTL) scenarios, we developed an efficient data layer processing and generation software, which we make publicly available. Using this tool, we created and released a fully-automated extension of the Dronescapes dataset, which is the largest of its kind in the literature, to the best of our knowledge. All the technical details, code and reproduction steps can be found on the project’s website (<a class="anchor anchor-primary" href="https://sites.google.com/view/dronescapes-dataset" target="_blank"><span class="anchor-text-container"><span class="anchor-text">https://sites.google.com/view/dronescapes-dataset</span>
<svg focusable="false" viewBox="0 0 8 8" height="20" aria-label="Opens in new window" class="icon icon-arrow-up-right-tiny arrow-external-link">
<path d="M1.12949 2.1072V1H7V6.85795H5.89111V2.90281L0.784057 8L0 7.21635L5.11902 2.1072H1.12949Z"></path>
</svg></span></a>).
Keywords:
Deep learning
Masked auto-encoders (MAE)
Multi-modal neural hypergraphs
Ensemble learning
Knowledge distillation
Aerial image understanding
Semi-supervised learning
Self-supervised learning
Unmanned aerial vehicles (UAVs)
Neural graph consensus
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

