Return
AMPEL workflows for LSST: Modular and reproducible real-time photometric classification
J
V
J
S
M
DOI:10.1051/0004-6361/202452481.png)
Abstract
En 中文
Context. Modern time-domain astronomical surveys produce high throughput data streams that require tools for processing and analysis. This will be critical for programs making full use of the alert stream from the Vera Rubin Observatory (VRO), where spectroscopic labels will only be available for a small subset of all transients. Aims. We introduce how the AMPEL toolset can work as a code-to-data platform for the development of efficient, reproducible and flexible workflows for real-time astronomical application. Methods. The Extended LSST Astronomical Time-series Classification Challenge (ELAsTiCC) v1 dataset contains a wide range of simulated astronomical transients, taking both the expected VRO noise profile and cadence into account. In this work, we introduce three different AMPEL channels constructed to highlight different uses of alert streams: to rapidly find infant transients (SNGuess), to provide unbiased transient samples for follow-up (FollowMe), and to deliver final transient classifications (FinalBet). These pipelines already contain placeholders for mechanisms that will be essential for the optimal usage of VRO alerts: combining different classifiers (built on boosted decision trees and deep neural networks), including host galaxy information, population priors, and sampling non-Gaussian photometric redshift distributions. Results. All three channels are already working at a high level: SNGuess correctly tags similar to 99% of all young supernovae, FollowMe illustrates how an unbiased subset of alerts can be selected for spectroscopic follow-up in the context of cosmological probes and FinalBet includes priors to achieve successful classifications for greater than or similar to 80% of all extragalactic transients. Conclusions. Advanced statistical tools, including machine learning, will be critical for the next decade of real-time astronomy. However, training these models are only initial steps as the scientific application in a real-time pipeline also relies on a long list of (conscious or unconscious) decisions. These include how data should be pre-filtered, probabilities combined, external information incorporated, and thresholds set for reactions. The fully functional workflows presented here are all public and can be used as starting points for any group wishing to optimize pipelines for their specific VRO science programs. AMPEL is designed to allow this to be done in accordance with FAIR principles: both software and results can be easily shared and results reproduced. The code-to-data environment ensures that models developed this way can be directly applied to the real-time LSST stream parsed by AMPEL.
Keywords:
methods: data analysis
methods: numerical
surveys
supernovae: general
stars: variables: general
Journal
IF:
5.8
Papers:
5.0W
Citations:
18.3W
