Skip to main content
QUICK REVIEW

[Paper Review] Improving Photometric Redshift Estimation for Cosmology with LSST using Bayesian Neural Networks

Evan Jones, Tuan Do|arXiv (Cornell University)|Jun 22, 2023
Advanced Statistical Methods and Models5 citations
TL;DR

This paper proposes a Bayesian Neural Network (BNN) for photometric redshift (photo-z) estimation in preparation for the LSST survey, using a Hyper Suprime-Cam data set with grizy photometry. The BNN provides accurate redshift predictions with well-calibrated uncertainty estimates, meeting 2/3 of LSST's photo-z science requirements and reducing catastrophic outliers by 66.8% when using a 0.8σ uncertainty cutoff.

ABSTRACT

We present results exploring the role that probabilistic deep learning models can play in cosmology from large-scale astronomical surveys through photometric redshift (photo-z) estimation. Photo-z uncertainty estimates are critical for the science goals of upcoming large-scale surveys such as LSST, however common machine learning methods typically provide only point estimates and lack uncertainties on predictions. We turn to Bayesian neural networks (BNNs) as a promising way to provide accurate predictions of redshift values with uncertainty estimates. We have compiled a galaxy data set from the Hyper Suprime-Cam Survey with grizy photometry, which is designed to be a smaller scale version of large surveys like LSST. We use this data set to investigate the performance of a neural network (NN) and a probabilistic BNN for photo-z estimation and evaluate their performance with respect to LSST photo-z science requirements. We also examine the utility of photo-z uncertainties as a means to reduce catastrophic outlier estimates. The BNN outputs the estimate in the form of a Gaussian probability distribution. We use the mean and standard deviation as the redshift estimate and uncertainty. We find that the BNN can produce accurate uncertainties. Using a coverage test, we find excellent agreement with expectation -- 67.2$\%$ of galaxies between $0 < 2.5$ have 1-$σ$ uncertainties that cover the spectroscopic value. We also include a comparison to alternative machine learning models using the same data. We find the BNN meets two out of three of the LSST photo-z science requirements in the range $0 < z < 2.5$.

Motivation & Objective

  • To develop a machine learning model that provides accurate photometric redshift predictions with well-calibrated uncertainties for large-scale cosmological surveys like LSST.
  • To evaluate whether Bayesian neural networks can meet the LSST photo-z science requirements, particularly regarding rms error, bias, and outlier control.
  • To assess the utility of photo-z uncertainties in identifying and reducing catastrophic outlier predictions.
  • To compare the BNN's performance against standard machine learning models (e.g., fully connected NN, random forest, XGBoost, SVM) on a representative, smaller-scale data set mimicking LSST.

Proposed method

  • The authors train a Bayesian Neural Network (BNN) on a newly compiled galaxy data set from the Hyper Suprime-Cam Survey with grizy photometry, simulating LSST-scale data.
  • The BNN outputs a Gaussian probability distribution for each galaxy, using the mean as the redshift estimate and the standard deviation as the uncertainty estimate.
  • The model is evaluated using LSST's science requirements, including rms error, bias, and 3σ catastrophic outlier fraction, with additional probabilistic metrics such as coverage and PIT histograms.
  • The BNN's uncertainty estimates are used to flag potential outliers by applying a zσ cutoff, where galaxies with predicted uncertainty > zσ are considered likely outliers.
  • The performance of the BNN is compared against a fully connected neural network, random forest, XGBoost, SVM, and other established photo-z methods on the same data set.
  • Statistical validation includes a coverage test to assess whether 68% of true spectroscopic redshifts fall within the 1σ uncertainty interval, and PIT (Probability Integral Transform) histogram analysis to evaluate calibration of the predictive distributions.

Experimental results

Research questions

  • RQ1Can a Bayesian Neural Network produce photometric redshift estimates that meet the LSST's scientific requirements for accuracy and uncertainty calibration?
  • RQ2How well do the uncertainty estimates from the BNN align with the true redshift distribution, as measured by coverage and PIT histogram analysis?
  • RQ3To what extent can photo-z uncertainties be used to identify and reduce catastrophic outlier predictions without sacrificing too many reliable estimates?
  • RQ4How does the BNN's performance compare to standard machine learning models (e.g., NN, RF, XGBoost) in terms of rms error, bias, and outlier fraction on a representative data set?

Key findings

  • The BNN achieves 67.2% coverage of true spectroscopic redshifts within their 1σ uncertainty interval, indicating well-calibrated uncertainty estimates.
  • The BNN meets two out of three LSST photo-z science requirements in the redshift range 0 < z < 2.5: it achieves an rms error below 0.2 and a bias below 0.003.
  • Using a zσ cutoff of 0.8, the BNN reduces the number of catastrophic outliers by 66.8% and total outliers by 25.0%, at the cost of removing only 3.2% of non-outlier galaxies.
  • The BNN outperforms other models, including fully connected neural networks, random forests, XGBoost, and SVMs, in terms of accuracy and uncertainty calibration on the same data set.
  • The PIT histogram shows that the BNN's predictive distributions are well-calibrated, with no significant over- or under-coverage.
  • The study identifies potential biases in the training data, such as underrepresentation of high-redshift galaxies and a bias toward luminous galaxies, which may affect performance at z > 1.5.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.