[Paper Review] Gaia Data Release 2: Summary of the variability processing & analysis results
This paper summarizes the variability processing and results from Gaia Data Release 2 (DR2), which identifies over 550,000 variable star candidates across the sky using three-band photometry (G, GBP, GRP) from the first 22 months of Gaia operations. It presents a probabilistic, automated classification pipeline that yields 228,904 RR Lyrae stars, 11,438 Cepheids, 151,761 long-period variables, 147,535 rotation-modulated stars, 8,882 δ Scuti/SX Phoenicis stars, and 3,018 short-timescale variables, with approximately half of these being newly identified sources.
The Gaia Data Release 2 (DR2): we summarise the processing and results of the identification of variable source candidates of RR Lyrae stars, Cepheids, long period variables (LPVs), rotation modulation (BY Dra-type) stars, delta Scuti & SX Phoenicis stars, and short-timescale variables. In this release we aim to provide useful but not necessarily complete samples of candidates. The processed Gaia data consist of the G, BP, and RP photometry during the first 22 months of operations as well as positions and parallaxes. Various methods from classical statistics, data mining and time series analysis were applied and tailored to the specific properties of Gaia data, as well as various visualisation tools. The DR2 variability release contains: 228'904 RR Lyrae stars, 11'438 Cepheids, 151'761 LPVs, 147'535 stars with rotation modulation, 8'882 delta Scuti & SX Phoenicis stars, and 3'018 short-timescale variables. These results are distributed over a classification and various Specific Object Studies (SOS) tables in the Gaia archive, along with the three-band time series and associated statistics for the underlying 550'737 unique sources. We estimate that about half of them are newly identified variables. The variability type completeness varies strongly as function of sky position due to the non-uniform sky coverage and intermediate calibration level of this data. The probabilistic and automated nature of this work implies certain completeness and contamination rates which are quantified so that users can anticipate their effects. This means that even well-known variable sources can be missed or misidentified in the published data. The DR2 variability release only represents a small subset of the processed data. Future releases will include more variable sources and data products; however, DR2 shows the (already) very high quality of the data and great promise for variability studies.
Motivation & Objective
- To process and classify variable stars in Gaia Data Release 2 using multi-epoch photometry from the first 22 months of Gaia operations.
- To provide a comprehensive, albeit incomplete, sample of variable star candidates across the entire sky, focusing on high-amplitude pulsators and rotation-modulated stars.
- To quantify completeness and contamination rates for each variability class due to non-uniform sky coverage and calibration limitations.
- To deliver time-series photometry, classification results, and Specific Object Studies (SOS) tables for the astronomical community.
- To lay the foundation for future Gaia data releases by demonstrating the high-quality potential of Gaia's photometric survey for variability science.
Proposed method
- Applied time-series analysis, classical statistics, and data mining techniques tailored to Gaia's photometric data characteristics.
- Used G, GBP, and GRP band photometry from field-of-view averaged transit measurements over 22 months (July 2014–May 2016).
- Implemented multiple filtering thresholds (≥2, ≥12, ≥20 FoV transits) to assess relative completeness and reduce noise.
- Employed supervised classification pipelines (geq2, geq12, geq20 paths) to assign variability types, with cross-validation via Specific Object Studies (SOS) tables.
- Utilized probabilistic classification to estimate contamination and completeness levels, with results quantified in Tables 2 and 3.
- Cross-matched results with external catalogues (e.g., OGLE, K2, Kepler) to validate classifications and assess reliability.
Experimental results
Research questions
- RQ1What is the distribution and classification of variable stars in Gaia DR2 across different variability types?
- RQ2How does non-uniform sky coverage and calibration level affect the completeness and reliability of variable star detection in Gaia DR2?
- RQ3To what extent are known variable stars missed or misclassified in the DR2 variability processing due to data limitations?
- RQ4How do the classification results from automated pipelines compare with ground-truth validation from external missions like Kepler and K2?
- RQ5What is the estimated number of newly identified variable stars in DR2, and how do contamination and completeness vary across different classes?
Key findings
- Gaia DR2 identifies 228,904 RR Lyrae stars, 11,438 Cepheids, 151,761 long-period variables (LPVs), 147,535 stars with rotation modulation, 8,882 δ Scuti and SX Phoenicis stars, and 3,018 short-timescale variables.
- Approximately half of the 550,737 variable sources released in DR2 are estimated to be newly identified, based on cross-matches with external catalogues.
- Relative completeness is 80% for sources with ≥12 FoV transits and 51% for those with ≥20 FoV transits, indicating significant sky coverage gaps.
- Contamination and completeness vary strongly by sky position due to non-uniform scanning law and calibration state, with relative completeness ranging from 50–70% for ≥20 transits.
- Overlaps between classes exist: 72 short-timescale objects overlap with RR Lyrae, 5 with Cepheids, and 3 with rotation modulation, reflecting overlapping physical definitions.
- In 618 cases, the SOS module reclassified sources from RR Lyrae to Cepheid, while only 77 were reclassified from Cepheid to RR Lyrae, indicating asymmetry in classification reliability.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.