Skip to main content
QUICK REVIEW

[论文解读] Machine Learning and Materials Informatics: Recent Applications and Prospects

Rampi Ramprasad, Rohit Batra|arXiv (Cornell University)|Jul 23, 2017
Machine Learning in Materials Science参考文献 86被引用 9
一句话总结

本文综述了机器学习与材料信息学领域的最新进展,重点介绍了基于数据驱动的代理模型,这些模型利用数值指纹和学习到的映射关系来预测材料性能。研究强调了在高通量筛选、逆向设计以及不确定性感知预测中的应用,其关键贡献在于加速材料发现进程,并实现对复杂、难以计算的性质的预测。

ABSTRACT

Propelled partly by the Materials Genome Initiative, and partly by the algorithmic developments and the resounding successes of data-driven efforts in other domains, informatics strategies are beginning to take shape within materials science. These approaches lead to surrogate machine learning models that enable rapid predictions based purely on past data rather than by direct experimentation or by computations/simulations in which fundamental equations are explicitly solved. Data-centric informatics methods are becoming useful to determine material properties that are hard to measure or compute using traditional methods--due to the cost, time or effort involved--but for which reliable data either already exists or can be generated for at least a subset of the critical cases. Predictions are typically interpolative, involving fingerprinting a material numerically first, and then following a mapping (established via a learning algorithm) between the fingerprint and the property of interest. Fingerprints may be of many types and scales, as dictated by the application domain and needs. Predictions may also be extrapolative--extending into new materials spaces--provided prediction uncertainties are properly taken into account. This article attempts to provide an overview of some of the recent successful data-driven "materials informatics" strategies undertaken in the last decade, and identifies some challenges the community is facing and those that should be overcome in the near future.

研究动机与目标

  • 综合评估过去十年机器学习在材料信息学中的最新应用。
  • 识别在数据质量、泛化能力及不确定性量化方面面临的关键挑战,以实现可靠预测。
  • 探讨利用机器学习实现目标性能材料逆向设计的可行性与局限性。
  • 为研究人员提供在何时以及如何在材料科学问题中有效应用机器学习的指导。
  • 推动机器学习与传统物理驱动方法的融合,以实现稳健、高效的材料发现。

提出的方法

  • 利用数值指纹(例如元素、结构、电子性质)将材料表示在高维空间中。
  • 采用监督式机器学习算法,学习材料指纹与目标性能之间的映射关系。
  • 应用代理模型,实现在无需求解基本方程的情况下的快速插值预测。
  • 整合不确定性量化,以评估预测的可靠性,尤其是在外推区域。
  • 使用约束优化和进化算法(例如遗传算法),通过将目标性能反向映射回材料组成,实现逆向设计。
  • 利用高通量模拟与实验生成可靠的训练数据,用于模型训练与验证。

实验结果

研究问题

  • RQ1在哪些类型的材料问题中,机器学习比传统物理建模更为有效?
  • RQ2如何使机器学习模型在已知数据分布之外的外推区域中保持鲁棒性?
  • RQ3在材料科学中利用机器学习实现可靠逆向设计面临哪些关键挑战?
  • RQ4如何量化机器学习预测中的不确定性,并用于指导实验或计算的后续工作?
  • RQ5研究人员应依据哪些标准判断某一材料问题是否适合采用数据驱动方法?

主要发现

  • 机器学习代理模型可实现对复杂材料性能(如玻璃化转变温度、介电损耗和机械强度)的快速预测,而无需直接计算或测量。
  • 数据驱动方法已成功从大规模数据集中识别出经验规律(例如胡姆-罗瑟里型趋势)和材料行为的相关性。
  • 指纹表示与学习映射的结合,使得在高维化学与构型空间中进行高效探索成为可能。
  • 不确定性量化对于可靠外推至关重要,并且正越来越多地被整合到预测框架中,以指导自适应学习。
  • 尽管逆向设计仍具挑战性,但约束优化与进化算法在生成目标性能候选材料方面展现出巨大潜力。
  • 机器学习力场已发展为一种强大工具,可通过替代昂贵的量子力学计算,显著加速分子动力学模拟。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。