[论文解读] Accurate Prediction of Experimental Band Gaps from Large Language Model-Based Data Extraction
本文提出了一种基于大语言模型(LLM)的流程,用于从科学文献中提取实验带隙数据,其错误率显著低于以往的自动化方法。通过仅筛选纯的、单晶块体材料并基于由此产生的高质量数据集进行训练,作者将带隙预测的平均绝对误差相比现有人工整理的数据库降低了19%。
Machine learning is transforming materials discovery by providing rapid predictions of material properties, which enables large-scale screening for target materials. However, such models require training data. While automated data extraction from scientific literature has potential, current auto-generated datasets often lack sufficient accuracy and critical structural and processing details of materials that influence the properties. Using band gap as an example, we demonstrate Large language model (LLM)-prompt-based extraction yields an order of magnitude lower error rate. Combined with additional prompts to select a subset of experimentally measured properties from pure, single-crystalline bulk materials, this results in an automatically extracted dataset that's larger and more diverse than the largest existing human-curated database of experimental band gaps. Compared to the existing human-curated database, we show the model trained on our extracted database achieves a 19% reduction in the mean absolute error of predicted band gaps. Finally, we demonstrate that LLMs are able to train models predicting band gap on the extracted data, achieving an automated pipeline of data extraction to materials property prediction.
研究动机与目标
- 解决材料科学中机器学习模型训练数据稀缺且不准确的问题。
- 提高从科学文献中自动化提取材料性质数据的可靠性。
- 开发一种可扩展、低错误率的流程,用于从全文论文中提取实验带隙数据。
- 证明基于LLM提取的数据可训练出优于基于现有整理数据库训练的模型。
- 建立从文献到性质预测的端到端自动化工作流程。
提出的方法
- 采用少样本提示(few-shot prompting)结合大语言模型,从科学论文中提取带隙数值。
- 应用额外的LLM提示,仅筛选实验测量的、纯的、单晶块体材料。
- 从经过筛选的LLM提取数据中构建一个大规模、多样化的实验带隙数据集。
- 在LLM提取的数据集上训练机器学习模型,用于带隙预测。
- 将模型性能与现有最大的人工整理的实验带隙数据库进行验证。
- 采用多步LLM提示策略,以提升数据质量并减少噪声。
实验结果
研究问题
- RQ1基于LLM的数据提取是否能在带隙提取中实现比现有自动化方法更低的错误率?
- RQ2将LLM提取的数据仅限于纯的、单晶块体材料是否能提升数据质量与预测性能?
- RQ3在LLM提取数据上训练的机器学习模型是否能优于在人工整理数据库上训练的模型?
- RQ4LLM提取数据集的规模与多样性与现有整理数据库相比如何?
- RQ5从科学文献到材料性质预测的端到端自动化流程是否可行?
主要发现
- 基于LLM的提取方法相比以往的自动化数据提取技术,错误率降低了整整一个数量级。
- 所生成的数据集比现有最大的人工整理的实验带隙数据库更大、更具多样性。
- 在LLM提取数据集上训练的模型,相比在人工整理数据库上训练的模型,带隙预测的平均绝对误差降低了19%。
- LLM流程成功以高保真度识别并筛选出纯的、单晶块体材料的实验测量带隙。
- 本研究展示了使用LLM实现从科学文献到准确材料性质预测的完全自动化、可扩展的端到端流程。
- LLM提取数据集带来的性能提升表明,数据质量与代表性对模型准确性至关重要。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。