Skip to main content
QUICK REVIEW

[논문 리뷰] LLM4Mat-Bench: Benchmarking Large Language Models for Materials Property Prediction

Andre Niyongabo Rubungo, Kangming Li|arXiv (Cornell University)|2024. 10. 31.
Machine Learning in Materials Science인용 수 6
한 줄 요약

LLM4Mat-Bench는 다양한 LLM이 구성, CIF, 또는 텍스트 설명을 사용하여 결정 물질의 특성을 예측하는 능력을 대규모 벤치마크로 평가하며, 물질 특성 예측에서 일반 목적 LLM보다 작업 특화 모델의 우수성을 강조한다.

ABSTRACT

Large language models (LLMs) are increasingly being used in materials science. However, little attention has been given to benchmarking and standardized evaluation for LLM-based materials property prediction, which hinders progress. We present LLM4Mat-Bench, the largest benchmark to date for evaluating the performance of LLMs in predicting the properties of crystalline materials. LLM4Mat-Bench contains about 1.9M crystal structures in total, collected from 10 publicly available materials data sources, and 45 distinct properties. LLM4Mat-Bench features different input modalities: crystal composition, CIF, and crystal text description, with 4.7M, 615.5M, and 3.1B tokens in total for each modality, respectively. We use LLM4Mat-Bench to fine-tune models with different sizes, including LLM-Prop and MatBERT, and provide zero-shot and few-shot prompts to evaluate the property prediction capabilities of LLM-chat-like models, including Llama, Gemma, and Mistral. The results highlight the challenges of general-purpose LLMs in materials science and the need for task-specific predictive models and task-specific instruction-tuned LLMs in materials property prediction.

연구 동기 및 목표

  • 물질 특성 예측에서 LLM을 평가하기 위한 표준화된 벤치마크의 필요성을 촉구한다.
  • 다양한 데이터 소스, 모달리티 및 특성을 갖춘 포괄적이고 다양한 벤치마크(LLM4Mat-Bench)를 구축한다.
  • 작업 특화 예측 모델에서 일반 목적 LLM에 이르는 다양한 모델을 평가하여 강점과 한계를 파악한다.

제안 방법

  • 중복을 제거한 후 10개의 데이터 소스로부터 약 ~1.9M 결정 구조를 수집하여 1,978,985개의 구성–구조–설명 쌍으로 구성한다.
  • Robocrystallographer를 사용해 결정 구조 설명을 결정론적으로 생성하여 데이터 오염이 없는 텍스트 기반 입력 모달리티를 만든다.
  • LLM-Prop, MatBERT, Llama, Gemma, Mistral, CGCNN를 기준선으로 삼아 다수의 모델 계열에 걸쳐 세 가지 물질 표현(Composition, CIF, Description)을 평가한다.
  • 작고 작업 특화된 모델(LLM-Prop, MatBERT)을 미세조정하고, 더 큰 채팅형 LLM의 제로샷 및 페어샷 프롬프트와 비교한다.
  • 재현 가능한 비교를 가능하게 하기 위해 고정된 학습/검증/테스트 분할과 표준 지표(MAD:MAE 회귀, AUC 분류)를 사용한다.

실험 결과

연구 질문

  • RQ1다양한 데이터 소스와 입력 모달리티에 걸쳐 LLM이 물질 특성 예측에 효과적으로 사용될 수 있는가?
  • RQ2작업 특화된 작고 LLM들이 물질 특성 예측에서 일반 목적의 대화형 LLM보다 우수한가?
  • RQ3어떤 입력 표현(Composition, CIF, Description)이 LLM 기반 모델에서 가장 높은 예측 성능을 낳는가?
  • RQ4이 분야에서 채팅형 LLM의 프롬프트 기반 제로샷 및 페어샷 평가가 미세 조정된 예측 모델과 어떻게 비교되는가?

주요 결과

  • 작업 특화된 소형 예측 LLM(LLM-Prop 및 MatBERT)이 회귀 및 분류 과제에서 일반 목적의 채팅형 LLM보다 우수하다.
  • 설명 기반 입력이 CIF 또는 구성 입력보다 LLM 기반 특성 예측기에 일반적으로 더 나은 성능을 낳는다.
  • 더 진보된 대형 생성형 LLM은 개선이 제한적이며 종종 물질 특성에 대해 잘못된 출력이나 허구를 생성한다.
  • 에너지적 특성은 데이터셋 전반에 걸쳐 다른 특성 유형보다 더 정확하게 예측된다.
  • MP 데이터에 대한 미세 조정은 효과적일 수 있지만 데이터셋과 특성에 따라 이득이 다르고, 일반 LLM은 탁월하려면 작업 특화 튜닝이 필요하다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.