Skip to main content
QUICK REVIEW

[Paper Review] Using a Neural Network Classifier to Select Galaxies with the Most Accurate Photometric Redshifts

Adam Broussard, Eric Gawiser|arXiv (Cornell University)|Aug 30, 2021
Galaxies: Formation, Evolution, PhenomenaPhysics and Astronomy74 references3 citations
TL;DR

This paper proposes a custom neural network classifier (NNC) to select galaxies with the most accurate photometric redshifts from large surveys like LSST. By training on TPZ-fitted redshifts and photometric data, the NNC reduces outlier fractions and scatter (σz) beyond what is possible using standard photo-z uncertainties, achieving a 35% lower outlier rate and 23% lower σz when selecting the top third of galaxies compared to TPZ-only selection.

ABSTRACT

The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) will produce several billion photometric redshifts (photo-$z$'s), enabling cosmological analyses to select a subset of galaxies with the most accurate photo-$z$. We perform initial redshift fits on Subaru Strategic Program galaxies with deep $grizy$ photometry using Trees for Photo-Z (TPZ) before applying a custom neural network classifier (NNC) tuned to select galaxies with $(z_\mathrm{phot} - z_\mathrm{spec})/(1+z_\mathrm{spec}) < 0.10$. We consider four cases of training and test sets ranging from an idealized case to using data augmentation to increase the representation of dim galaxies in the training set. Selections made using the NNC yield significant further improvements in outlier fraction and photo-$z$ scatter ($\sigma_z$) over those made with typical photo-$z$ uncertainties. As an example, when selecting the best third of the galaxy sample, the NNC achieves a 35% improvement in outlier rate and a 23% improvement in $\sigma_z$ compared to using uncertainties from TPZ. For cosmology and galaxy evolution studies, this method can be tuned to retain a particular sample size or to achieve a desired photo-$z$ accuracy; our results show that it is possible to retain more than a third of an LSST-like galaxy sample while reducing $\sigma_z$ by a factor of two compared to the full sample, with one-fifth as many photo-$z$ outliers. For surveys like LSST that are not limited by shot noise, this method enables a larger number of tomographic redshift bins and hence a significant increase in the total signal-to-noise of galaxy angular power spectra.

Motivation & Objective

  • To improve photometric redshift accuracy for large galaxy surveys such as LSST by selecting galaxies with the most reliable redshift estimates.
  • To address the challenge of high outlier fractions and scatter in photometric redshifts, which degrade cosmological clustering measurements.
  • To develop a post-processing method that enhances existing photo-z codes without requiring retraining or template selection.
  • To demonstrate that data augmentation and neural network filtering can significantly improve sample quality while retaining a large fraction of the original galaxy sample.

Proposed method

  • A custom neural network classifier (NNC) is trained to identify galaxies with (zphot − zspec)/(1 + zspec) < 0.10, using TPZ-fitted redshifts and photometric features as input.
  • The NNC is trained on four training/test set configurations, including data augmentation to improve representation of dim galaxies.
  • The method uses a four-layer feedforward neural network with [100, 200, 100, 50] neurons and a sigmoid output to assign a confidence score for accurate redshift selection.
  • The NNC is applied as an afterburner to existing photo-z codes, such as TPZ and BPZ, to refine selection without altering the original fitting process.
  • A neural network regressor (NNR) is tested to correct systematic errors in TPZ outputs before classification, though it provides no significant improvement in this study.
  • Performance is evaluated using metrics including outlier fraction, NMAD, and σz across varying sample fractions (fsample) from 1% to 100%.

Experimental results

Research questions

  • RQ1Can a neural network classifier significantly reduce the outlier fraction and scatter (σz) in photometric redshift estimates beyond what is achievable using standard uncertainty-based cuts?
  • RQ2How does data augmentation, particularly for dim galaxies, affect the performance of the NNC in selecting high-accuracy photo-z galaxies?
  • RQ3To what extent can the NNC improve photo-z accuracy when applied to different photo-z codes, such as TPZ and BPZ?
  • RQ4Does the addition of a pre-classification neural network regressor (NNR) to correct systematic errors in TPZ outputs lead to measurable improvements in selection performance?
  • RQ5What is the optimal training sample size for the NNC to maximize selection performance in terms of outlier rejection and σz reduction?

Key findings

  • When selecting the top third of galaxies, the NNC reduces the outlier fraction by 35% and σz by 23% compared to selection based on TPZ uncertainties.
  • The NNC achieves a 35% improvement in outlier rate and 23% improvement in σz when selecting the best third of the sample, compared to using TPZ-reported uncertainties.
  • Training on approximately 50,000 galaxies (30% of the training set) is sufficient to maximize NNC performance for outlier rejection in the Match → COSMOS2015 case.
  • The NMAD of photo-z errors shows monotonic, incremental improvement with increasing training sample size, suggesting continued gains with larger datasets.
  • The NNC applied to BPZ-fitted redshifts achieves similar outlier and σz performance to TPZ-based selection, though with a constant NMAD offset of ~0.008.
  • The addition of a neural network regressor (NNR) to correct TPZ systematic errors did not yield meaningful improvements and was therefore excluded from the final pipeline.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.