Skip to main content
QUICK REVIEW

[Paper Review] Where's Swimmy?: Mining unique color features buried in galaxies by deep anomaly detection using Subaru Hyper Suprime-Cam data

Takumi S. Tanaka, Rhythm Shimakawa|arXiv (Cornell University)|Oct 11, 2021
Galaxies: Formation, Evolution, PhenomenaPhysics and Astronomy141 references9 citations
TL;DR

This paper introduces the Swimmy survey, a deep anomaly detection framework using autoencoders on Subaru Hyper Suprime-Cam multicolor imaging data to identify rare, unique galaxies without labeled training data. It successfully recovered 60–70% of known quasars and 60% of extreme emission-line galaxies (XELGs) as outliers, demonstrating that unsupervised anomaly detection can efficiently uncover rare and potentially novel astrophysical phenomena in large astronomical datasets.

ABSTRACT

We present the Swimmy (Subaru WIde-field Machine-learning anoMalY) survey program, a deep-learning-based search for unique sources using multicolored ($grizy$) imaging data from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). This program aims to detect unexpected, novel, and rare populations and phenomena, by utilizing the deep imaging data acquired from the wide-field coverage of the HSC-SSP. This article, as the first paper in the Swimmy series, describes an anomaly detection technique to select unique populations as "outliers" from the data-set. The model was tested with known extreme emission-line galaxies (XELGs) and quasars, which consequently confirmed that the proposed method successfully selected 60-70% of the quasars and 60% of the XELGs without labeled training data. In reference to the spectral information of local galaxies at $z=$0.05-0.2 obtained from the Sloan Digital Sky Survey, we investigated the physical properties of the selected anomalies and compared them based on the significance of their outlier values. The results revealed that XELGs constitute notable fractions of the most anomalous galaxies, and certain galaxies manifest unique morphological features. In summary, a deep anomaly detection is an effective tool that can search rare objects, and ultimately, unknown unknowns with large data-sets. Further development of the proposed model and selection process can promote the practical applications required to achieve specific scientific goals.

Motivation & Objective

  • To develop an unsupervised anomaly detection method for identifying rare and unique galaxies in large-scale imaging surveys.
  • To detect 'unknown unknowns'—previously undetected populations or phenomena—by identifying extreme outliers in galaxy color and morphology.
  • To validate the method using known extreme sources like quasars and XELGs, ensuring it can recover known rare objects without prior labeling.
  • To investigate the physical properties of detected anomalies using multi-wavelength archival data, linking outlier scores to astrophysical significance.

Proposed method

  • Employs a deep autoencoder neural network to learn a low-dimensional latent representation of galaxy images from grizy-band HSC-SSP data.
  • Uses reconstruction error as the anomaly score: higher error indicates greater deviation from typical galaxy features.
  • Applies a normalized anomaly score (Sanom) based on the z-score of reconstruction error relative to the training sample.
  • Trains the autoencoder on a representative sample of typical galaxies, then identifies outliers by ranking sources by Sanom.
  • Performs model selection via repeated training (30 runs) with fixed hyperparameters (d=8, rgauss=0.02), selecting the model with highest selection rate for known quasars and XELGs.
  • Validates results by cross-referencing top anomalies with spectroscopic data (e.g., SDSS DR15) and examining residuals for artifacts.

Experimental results

Research questions

  • RQ1Can deep anomaly detection identify rare and extreme galaxies without prior labeling or templates?
  • RQ2How effective is the anomaly detection method in recovering known extreme sources such as quasars and XELGs?
  • RQ3What physical properties distinguish the most anomalous galaxies identified by the model?
  • RQ4To what extent do artifacts or data reduction errors contaminate the anomaly candidate list?
  • RQ5Can the method be extended to color-selected samples in the local universe (z ≈ 0.05–0.2) for broader discovery potential?

Key findings

  • The model successfully recovered 60–70% of known DR16Q quasars and 60% of XELG samples as top anomalies, demonstrating strong detection capability without labeled training data.
  • Extreme emission-line galaxies (XELGs) constituted a significant fraction of the most anomalous galaxies, indicating their distinct SEDs are effectively captured by the anomaly score.
  • The top 0.0465% of anomalies in a color-selected local sample (z ≈ 0.05–0.2) included numerous blue, green, and purple compact sources, suggesting potential new XELG candidates.
  • A significant number of false positives were traced to artifacts, particularly in the r-band, caused by flux normalization errors or zero-point miscalibrations, highlighting the need for data quality control.
  • The method remains robust across multiple training runs, with consistent trends despite stochasticity, and model selection based on selection rate for known sources proved effective.
  • The approach enables the discovery of 'unknown unknowns' by identifying sources that deviate significantly from typical galaxy SEDs and morphologies, even when no prior examples exist.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.