[Paper Review] The CAMELS project: public data release
The CAMELS project releases 4,233 cosmological hydrodynamic simulations, 2,049 N-body simulations, and 2,184 hydrodynamic simulations across two distinct codes (IllustrisTNG and SIMBA), along with derived data products such as galaxy, halo, and void catalogs, power spectra, and Lyman-α spectra. This comprehensive, publicly available dataset enables machine learning applications in cosmology and astrophysics, supporting studies from galaxy mass constraints to cross-field information extraction.
The Cosmology and Astrophysics with MachinE Learning Simulations (CAMELS) project was developed to combine cosmology with astrophysics through thousands of cosmological hydrodynamic simulations and machine learning. CAMELS contains 4,233 cosmological simulations, 2,049 N-body and 2,184 state-of-the-art hydrodynamic simulations that sample a vast volume in parameter space. In this paper we present the CAMELS public data release, describing the characteristics of the CAMELS simulations and a variety of data products generated from them, including halo, subhalo, galaxy, and void catalogues, power spectra, bispectra, Lyman-$α$ spectra, probability distribution functions, halo radial profiles, and X-rays photon lists. We also release over one thousand catalogues that contain billions of galaxies from CAMELS-SAM: a large collection of N-body simulations that have been combined with the Santa Cruz Semi-Analytic Model. We release all the data, comprising more than 350 terabytes and containing 143,922 snapshots, millions of halos, galaxies and summary statistics. We provide further technical details on how to access, download, read, and process the data at \url{https://camels.readthedocs.io}.
Motivation & Objective
- To create a large, publicly accessible dataset of cosmological simulations spanning diverse physical parameter space to bridge cosmology and astrophysics.
- To enable machine learning applications by providing high-fidelity hydrodynamic and N-body simulations with consistent metadata and derived products.
- To support the community with interactive, scalable, and efficient data access through Binder, Globus, and FlatHUB platforms.
- To address limitations in existing simulations by offering a broad parameter space coverage despite small individual simulation volumes.
- To facilitate robust, reproducible research by documenting all data products and ensuring long-term updates via centralized online documentation.
Proposed method
- Simulate 4,233 cosmological hydrodynamic simulations using two independent codes: IllustrisTNG and SIMBA, each with distinct subgrid physics models.
- Generate 2,049 N-body simulations as non-radiative counterparts to the hydrodynamic runs, preserving initial conditions and cosmological parameters.
- Produce derived data products including subhalo and galaxy catalogs (via Subfind and SAM), power spectra, bispectra, Lyman-α transmission spectra, and X-ray photon lists.
- Create CAMELS-SAM, a collection of over 1,000 semi-analytic galaxy catalogs derived from hundreds of N-body simulations using the Santa Cruz model.
- Organize data into structured, version-controlled directories with redshift-specific outputs and metadata, accessible via multiple platforms.
- Deploy interactive tools including a Binder environment for Jupyter-based exploration and Globus for high-throughput data transfer.
Experimental results
Research questions
- RQ1Can machine learning models trained on CAMELS data accurately infer cosmological parameters from galaxy and large-scale structure observables?
- RQ2To what extent can neural networks extract physical information from diverse astrophysical fields while marginalizing over subgrid physics variations?
- RQ3How well do simulated galaxy properties in CAMELS reproduce observed galaxy mass functions and clustering patterns?
- RQ4Can the CAMELS dataset constrain the masses of the Milky Way and Andromeda galaxies using artificial intelligence techniques?
- RQ5How do different hydrodynamic codes (IllustrisTNG vs. SIMBA) affect the statistical properties of galaxies and halos across the parameter space?
Key findings
- The CAMELS project provides 4,233 hydrodynamic simulations and 2,049 N-body simulations, covering a vast parameter space with two distinct subgrid physics implementations.
- Derived data products include galaxy, halo, and subhalo catalogs, power spectra, bispectra, Lyman-α spectra, radial profiles, and X-ray photon lists, all accessible via standardized interfaces.
- CAMELS-SAM delivers over 1,000 semi-analytic galaxy catalogs from N-body simulations, enabling comparison across different merger trees and model assumptions.
- The dataset is hosted across multiple platforms, including Binder for interactive analysis, Globus for high-speed transfer, and FlatHUB for efficient exploration of Subfind catalogs.
- The project enables the first AI-based constraints on the masses of the Milky Way and Andromeda galaxies using simulated observables.
- Neural networks trained on CAMELS data successfully extract information from vastly different physical fields while marginalizing over astrophysical uncertainties at the field level.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.