Skip to main content
QUICK REVIEW

[论文解读] Is Large Language Model All You Need to Predict the Synthesizability and Precursors of Crystal Structures?

Zhilong Song, Shuaihua Lu|arXiv (Cornell University)|Jul 9, 2024
Machine Learning in Materials Science被引用 4
一句话总结

该论文提出了晶体合成大语言模型(CSLLM)框架,利用微调的大语言模型从晶体结构数据中预测晶体可合成性、合成方法及前驱体。该框架在包含140,120个样本的数据集上进行训练,采用一种新颖的文本表示方法,其在可合成性预测中的准确率达到98.6%,相较于热力学稳定性和动力学稳定性筛选方法分别提升了106.1%和44.5%。

ABSTRACT

Accessing the synthesizability of crystal structures is pivotal for advancing the practical application of theoretical material structures designed by machine learning or high-throughput screening. However, a significant gap exists between the actual synthesizability and thermodynamic or kinetic stability, which is commonly used for screening theoretical structures for experiments. To address this, we develop the Crystal Synthesis Large Language Models (CSLLM) framework, which includes three LLMs for predicting the synthesizability, synthesis methods, and precursors. We create a comprehensive synthesizability dataset including 140,120 crystal structures and develop an efficient text representation method for crystal structures to fine-tune the LLMs. The Synthesizability LLM achieves a remarkable 98.6% accuracy, significantly outperforming traditional synthesizability screening based on thermodynamic and kinetic stability by 106.1% and 44.5%, respectively. The Methods LLM achieves a classification accuracy of 91.02%, and the Precursors LLM has an 80.2% success rate in predicting synthesis precursors. Furthermore, we develop a user-friendly graphical interface that enables automatic predictions of synthesizability and precursors from uploaded crystal structure files. Through these contributions, CSLLM bridges the gap between theoretical material design and experimental synthesis, paving the way for the rapid discovery of novel and synthesizable functional materials.

研究动机与目标

  • 弥合材料设计中理论晶体稳定性与实际实验可合成性之间的差距。
  • 克服基于热力学和动力学稳定性的传统筛选方法的局限性。
  • 开发一个统一框架,用于预测晶体结构的可合成性、合成方法及前驱体材料。
  • 为实验研究人员创建一个用户友好的界面,可直接从晶体文件预测合成可行性。
  • 弥合高通量理论材料设计与实际实验合成之间的鸿沟。

提出的方法

  • 开发一种新颖的文本表示方法,将晶体结构转换为适合大语言模型处理的序列,利用结构和元素信息。
  • 在经过筛选的140,120个晶体结构数据集上,微调三个不同的大语言模型(可合成性LLM、方法LLM、前驱体LLM)。
  • 训练可合成性LLM,基于历史合成数据判断晶体结构是否具有实验可合成性。
  • 训练方法LLM,预测目标晶体结构的可行合成路径(例如固相法、溶剂热法等)。
  • 训练前驱体LLM,识别合成目标晶体所需的化学前驱体,重点关注化学计量比和反应相容性。
  • 实现一个图形用户界面,可接受晶体结构文件(如CIF格式),并返回关于可合成性和前驱体的自动化预测结果。

实验结果

研究问题

  • RQ1大语言模型能否在超越热力学和动力学稳定性的基础上,有效预测晶体结构的可合成性?
  • RQ2大语言模型在多大程度上能泛化到多样化的晶体结构,以推荐可行的合成方法?
  • RQ3大语言模型在多大程度上能准确预测目标晶体相的化学上可行的前驱体组合?
  • RQ4统一的LLM框架能否在现实世界中的可合成性预测中超越基于稳定性的传统筛选方法?
  • RQ5在少样本或零样本设置下,基于文本的晶体结构表示方法能实现怎样的性能水平?

主要发现

  • 可合成性LLM的预测准确率达到98.6%,显著优于热力学稳定性筛选(提升106.1%)和动力学稳定性筛选(提升44.5%)。
  • 方法LLM在合成路径分类上达到91.02%的准确率,可实现对实验方案的可靠推荐。
  • 前驱体LLM在80.2%的案例中成功预测出有效的前驱体组合,展现出强大的化学推理能力。
  • 所提出的文本表示方法能有效将晶体结构信息编码为自然语言标记,支持高性能的大语言模型微调。
  • 所开发的图形用户界面可实现从晶体结构输入文件到预测结果的端到端处理,显著提升了实验研究人员的使用便利性。
  • 该框架成功弥合了理论材料设计与实验可行性之间的关键鸿沟,加速了新型功能材料的发现。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。