[Paper Review] Accurate Prediction of Experimental Band Gaps from Large Language Model-Based Data Extraction
This paper presents a large language model (LLM)-based pipeline for extracting experimental band gap data from scientific literature, achieving significantly lower error rates than prior automated methods. By filtering for pure, single-crystalline bulk materials and training on the resulting high-quality dataset, the authors reduce mean absolute error in band gap prediction by 19% compared to existing human-curated databases.
Machine learning is transforming materials discovery by providing rapid predictions of material properties, which enables large-scale screening for target materials. However, such models require training data. While automated data extraction from scientific literature has potential, current auto-generated datasets often lack sufficient accuracy and critical structural and processing details of materials that influence the properties. Using band gap as an example, we demonstrate Large language model (LLM)-prompt-based extraction yields an order of magnitude lower error rate. Combined with additional prompts to select a subset of experimentally measured properties from pure, single-crystalline bulk materials, this results in an automatically extracted dataset that's larger and more diverse than the largest existing human-curated database of experimental band gaps. Compared to the existing human-curated database, we show the model trained on our extracted database achieves a 19% reduction in the mean absolute error of predicted band gaps. Finally, we demonstrate that LLMs are able to train models predicting band gap on the extracted data, achieving an automated pipeline of data extraction to materials property prediction.
Motivation & Objective
- To address the scarcity and inaccuracy of training data for machine learning models in materials science.
- To improve the reliability of automated data extraction from scientific literature for materials properties.
- To develop a scalable, low-error pipeline for extracting experimental band gaps from full-text papers.
- To demonstrate that LLM-extracted data can train models that outperform those trained on existing curated databases.
- To establish an end-to-end automated workflow from literature to property prediction.
Proposed method
- Employing few-shot prompting with large language models to extract band gap values from scientific papers.
- Applying additional LLM prompts to filter for experimentally measured, pure, single-crystalline bulk materials only.
- Constructing a large, diverse dataset of experimental band gaps from the filtered LLM-extracted data.
- Training a machine learning model on the LLM-extracted dataset for band gap prediction.
- Validating model performance against the largest existing human-curated database of experimental band gaps.
- Using a multi-step LLM prompting strategy to enhance data quality and reduce noise.
Experimental results
Research questions
- RQ1Can LLM-based data extraction achieve lower error rates in band gap extraction than existing automated methods?
- RQ2Does filtering LLM-extracted data to include only pure, single-crystalline bulk materials improve data quality and predictive performance?
- RQ3Can a machine learning model trained on LLM-extracted data outperform models trained on human-curated databases?
- RQ4To what extent does the scale and diversity of the LLM-extracted dataset compare to existing curated databases?
- RQ5Is it feasible to create an end-to-end automated pipeline from scientific literature to materials property prediction?
Key findings
- The LLM-based extraction method achieved an order of magnitude lower error rate compared to prior automated data extraction techniques.
- The resulting dataset is larger and more diverse than the largest existing human-curated database of experimental band gaps.
- A model trained on the LLM-extracted dataset reduced mean absolute error in band gap prediction by 19% compared to models trained on the human-curated database.
- The LLM pipeline successfully identified and filtered for experimentally measured band gaps in pure, single-crystalline bulk materials with high fidelity.
- The study demonstrates a fully automated, scalable pipeline from scientific literature to accurate materials property prediction using LLMs.
- The performance gain from the LLM-extracted dataset indicates that data quality and representativeness are critical for model accuracy.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.