Skip to main content
QUICK REVIEW

[论文解读] Modeling assembly bias with machine learning and symbolic regression

Digvijay Wadekar, Francisco Villaescusa-Navarro|arXiv (Cornell University)|Nov 30, 2020
Image Processing and 3D Reconstruction参考文献 103被引用 15
一句话总结

本文提出一种机器学习框架,利用IllustrisTNG流体动力学模拟对晕中性氢(H i)含量的组装偏差进行建模,表明环境依赖性(超越晕质量)显著影响H i功率谱。通过结合符号回归与随机森林,作者推导出解析表达式,提升了模拟星系目录的准确性,当忽略环境依赖性时,可将k ≳ 0.05 h Mpc⁻¹处21-cm功率谱的低估程度减少超过10%。

ABSTRACT

Upcoming 21cm surveys will map the spatial distribution of cosmic neutral hydrogen (HI) over unprecedented volumes. Mock catalogues are needed to fully exploit the potential of these surveys. Standard techniques employed to create these mock catalogs, like Halo Occupation Distribution (HOD), rely on assumptions such as the baryonic properties of dark matter halos only depend on their masses. In this work, we use the state-of-the-art magneto-hydrodynamic simulation IllustrisTNG to show that the HI content of halos exhibits a strong dependence on their local environment. We then use machine learning techniques to show that this effect can be 1) modeled by these algorithms and 2) parametrized in the form of novel analytic equations. We provide physical explanations for this environmental effect and show that ignoring it leads to underprediction of the real-space 21-cm power spectrum at $k\gtrsim 0.05$ h/Mpc by $\gtrsim$10\%, which is larger than the expected precision from upcoming surveys on such large scales. Our methodology of combining numerical simulations with machine learning techniques is general, and opens a new direction at modeling and parametrizing the complex physics of assembly bias needed to generate accurate mocks for galaxy and line intensity mapping surveys.

研究动机与目标

  • 研究晕H i含量对除质量外的次级晕属性(特别是局域环境)的依赖性。
  • 开发一种机器学习方法,从高保真流体动力学模拟中捕捉复杂且非线性的组装偏差效应。
  • 利用符号回归从模拟数据中推导出可解释的解析表达式,实现大规模模拟星表生成中的高效应用。
  • 量化忽略环境组装偏差对21-cm功率谱的影响,表明标准HOD模型中存在显著的系统性误差。
  • 提供一种可推广的方法,通过模拟-机器学习集成,用于星系和强度映射巡天中组装偏差的建模。

提出的方法

  • 使用IllustrisTNG模拟提取晕属性,包括H i质量、质量、环境(δ₂.₅)、集中度、自旋参数以及子晕质量分数。
  • 应用随机森林回归,将实际H i质量与HOD预测H i质量的比值建模为次级晕属性的函数。
  • 采用符号回归(通过PySR实现),仅使用加法和乘法运算符推导H i质量比的可解释解析表达式,以避免过拟合。
  • 在全物理(FP)模拟数据上校准HOD模型,并在仅含暗物质的(DMO)N体模拟上测试其性能,以评估质量不匹配的影响。
  • 通过与真实H i功率谱对比,验证预测结果,比较仅质量HOD、机器学习增强模型与模拟的结果。
  • 通过选择关键的次级参数(如环境、子晕分数)减少高维输入空间,使符号回归可行。

实验结果

研究问题

  • RQ1暗物质晕的H i含量在超越质量之外,如何依赖于其局域环境?
  • RQ2如随机森林和符号回归等机器学习模型能否准确捕捉并参数化这种环境依赖性?
  • RQ3忽略环境组装偏差对大尺度巡天中21-cm功率谱的定量影响是什么?
  • RQ4符号回归能否生成简洁、可解释的解析方程,使其在建模组装偏差方面优于标准HOD模型?
  • RQ5当应用于晕质量不匹配的DMO模拟时,基于机器学习的模拟星表性能与标准HOD模型相比如何?

主要发现

  • IllustrisTNG中晕的H i含量对局域环境(δ₂.₅)表现出强烈依赖性,较高密度区域对应更高的H i质量。
  • 基于次级晕属性训练的随机森林模型,与仅质量HOD相比,在k ≳ 0.05 h Mpc⁻¹处将H i功率谱预测偏差减少最多约10%。
  • 符号回归成功推导出紧凑的解析表达式(如公式6),将H i质量比建模为环境与子晕分数的函数,提升了模拟星表的准确性。
  • 忽略环境组装偏差导致真实空间21-cm功率谱在k ≳ 0.05 h Mpc⁻¹处低估超过10%,超过未来巡天的预期精度。
  • 在全物理模拟上校准的HOD模型在应用于DMO模拟时性能显著下降,这是由于晕质量不匹配所致,凸显了引入环境感知校正的必要性。
  • 符号回归推导的方程与模拟数据高度一致,残差处于预期散射范围内,验证了其物理相关性与预测能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。