[論文レビュー] Automated Extraction of Multicomponent Alloy Data Using Large Language Models for Sustainable Design
論文は、HEA文献全体のテキストと表から合金データを抽出する、LLMベースの二段階パイプラインを開発し、持続可能な材料設計のための大規模データベースを構築し、3領域への適用を実証する。
The design of sustainable materials requires access to materials performance and sustainability data from literature corpus in an organized, structured and automated manner. Natural language processing approaches, particularly large language models (LLMs), have been explored for materials data extraction from the literature, yet often suffer from limited accuracy or narrow scope. In this work, an LLM-based pipeline is developed to accurately extract alloy-related information from both textual descriptions and tabular data across the literature on high-entropy (or multicomponent) alloys (HEA). Specifically two databases with 37,711 and 148,069 entries respectively are retrieved; one from the literature text, consisting of alloy composition, processing conditions, characterization methods, and reported properties, and other from the literature tables, consisting of property names, values, and units. The pipeline enhances materials-domain sensitivity through prompt engineering and retrieval-augmented generation and achieves F1-scores of 0.83 for textual extraction and 0.88 for tabular extraction, surpassing or matching existing approaches. Application of the pipeline to over 10,000 articles yields the largest publicly available multicomponent alloy database and reveals compositional and processing-property trends. The database is further employed for sustainability-aware materials selection in three application domains, i.e., lightweighting, soft magnetic, and corrosion-resistant, identifying multicomponent alloy candidates with more sustainable production while maintaining or exceeding benchmark performance. The pipeline developed can be easily generalized to other class of materials, and assist in development of comprehensive, accurate and usable databases for sustainable materials design.
研究の動機と目的
- 持続可能な材料設計のために、非構造化された文献を機械可読な構造データへ変換する必要性を動機づける。
- LLMsを用いて、さまざまな合金報告スタイルを横断するテキストと表データを扱える堅牢で一般化可能なデータ抽出パイプラインを開発する。
- テキストと表からなる2つの包括的なデータベースを作成し、下流の持続可能性を考慮した材料選択を可能にする。
- HEAsおよびそれ以外の領域での研究を支援する curated データベースを公開する。
提案手法
- 二段階抽出パイプライン: (i) 段落レベルのテキスト抽出で合金系、加工、表征、特性を把握;(ii) 表を用いた抽出で特性値、単位、条件を把握。
- クエリセット1(QS1)は、プロンプト設計、 Few-shotデモンストレーション、検索拡張生成(RAG)を用いて、要約と実験セクションから合金組成、加工、特性を同定。
- クエリセット2(QS2)は、表セルを354項目の curated マスタ特性語彙にマッピングし、標準化された特性名を最初に特定し、次に対応する値と条件を抽出する二段階のLLMアプローチを適用。
- マスタ特性語彙は、正規化と拡張された記号/名称セットを備えた DB1 から構築され、表マッピングの堅牢性を高める。
- 評価は、欠落、幻視、新規エントリを考慮した拡張混同行列フレームワークを用い、精度・再現率・F1のトレードオフを強調。
- パイプラインはコストと精度のバランスのためにGPT-4oおよびGPT-4o miniを選択し、RAGを用いてベクトルデータベースに埋め込まれた98件の専門家アノテーション例を用いたFew-shotデモを適用。

実験結果
リサーチクエスチョン
- RQ1LLMベースのパイプラインは、テキストと表の両方からHEA文献コーパス全体の合金組成、加工の詳細、特性を正確に抽出できるか。
- RQ2テキスト(QS1)と表(QS2)の抽出における専門家ベンチマークと比較した場合、達成可能な精度・再現率・F1はどの程度か。
- RQ3得られたデータベースはどの程度大きく、どれだけ実用的か、複数領域の持続可能性を考慮した材料選択にどの程度情報を提供できるか。
- RQ4多成分合金のLLMベース抽出の実践的課題と限界は何で、どのように緩和できるか。
主な発見
- 2つのデータベースを作成: テキスト由来の合金レコード37,711件(DB1)と表由来のレコード148,069件(DB2)、出典は10,829編の論文。
- QS1は専門家アノテーション付きレビュー data に対してF1スコア約0.83を達成。
- QS2は広範なテストセットでF1約0.88、機械的特性に焦点を当てたセットで0.96を達成。
- テキスト抽出の精度/再現率は、レビュー用データでそれぞれ0.81と0.86、表抽出の精度/再現率はそれぞれ0.98と0.81。
- 構築されたデータベースは、軽量構造材、軟磁性、耐腐食分野における持続可能性を意識した選択を可能とし、性能を犠牲にすることなく多成分合金の持続可能性を改善。
- Alloy Tattvasarプラットフォームは、キュレーション済みデータベースをコミュニティが再利用できるよう公開に提供。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。