Skip to main content
QUICK REVIEW

[논문 리뷰] LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation

Yuan Chiang, Hsieh, Elvis|arXiv (Cornell University)|2024. 01. 30.
Machine Learning in Materials Science인용 수 21
한 줄 요약

LLaMP는 계층적 ReAct 에이전트를 활용하여 Materials Project의 고충실도 데이터에 기반해 LLM을 Ground하는 다중모달 회수-생성 프레임워크로, 파인튜닝 없이도 허위 진술을 줄입니다.

ABSTRACT

Reducing hallucination of Large Language Models (LLMs) is imperative for use in the sciences, where reliability and reproducibility are crucial. However, LLMs inherently lack long-term memory, making it a nontrivial, ad hoc, and inevitably biased task to fine-tune them on domain-specific literature and data. Here we introduce LLaMP, a multimodal retrieval-augmented generation (RAG) framework of hierarchical reasoning-and-acting (ReAct) agents that can dynamically and recursively interact with computational and experimental data on Materials Project (MP) and run atomistic simulations via high-throughput workflow interface. Without fine-tuning, LLaMP demonstrates strong tool usage ability to comprehend and integrate various modalities of materials science concepts, fetch relevant data stores on the fly, process higher-order data (such as crystal structure and elastic tensor), and streamline complex tasks in computational materials and chemistry. We propose a simple metric combining uncertainty and confidence estimates to evaluate the self-consistency of responses by LLaMP and vanilla LLMs. Our benchmark shows that LLaMP effectively mitigates the intrinsic bias in LLMs, counteracting the errors on bulk moduli, electronic bandgaps, and formation energies that seem to derive from mixed data sources. We also demonstrate LLaMP's capability to edit crystal structures and run annealing molecular dynamics simulations using pre-trained machine-learning force fields. The framework offers an intuitive and nearly hallucination-free approach to exploring and scaling materials informatics, and establishes a pathway for knowledge distillation and fine-tuning other language models. Code and live demo are available at https://github.com/chiang-yuan/llamp

연구 동기 및 목표

  • 과재현성 및 최신 데이터 접근을 위한 과학 분야의 신뢰 가능한 기억 기능을 갖춘 LLM의 필요성 동기를 제시합니다.
  • 다중모달 데이터 소스를 활용한 회수-강화된 기억 기반 프레임워크(LLaMP)를 재료 과학 정보학에 적용합니다.
  • 파인튜닝 없이 재료 데이터의 수집, 처리, 합성을 수행하는 위계 기반 ReAct 에이전트 오케스트레이션을 시연합니다.
  • 재료 특성 및 결정 구조에 대한 GPT-3.5 고유 지식과 LLaMP를 비교하여 허위 진술 감소를 정량화합니다.

제안 방법

  • LLMs를 Materials Project, arXiv, Wikipedia와 API 호출을 통해 연결하는 다중모달 회수-생성(RAG) 프레임워크를 구현합니다.
  • 상위 수준 에이전트가 도구 키트와 데이터 저장소를 갖춘 하위 수준 특화 에이전트를 조정하는 위계적 다중 에이전트 ReAct 계획을 사용합니다.
  • 추론과 작용을 고충실도 데이터로 접지하기 위해 API 스키마와 기억 버퍼를 도입합니다.
  • 계획된 작업 분해를 통해 고차원 데이터 처리(텐서, 결정 구조) 및 다중모달 추론을 시演합니다.
  • 형성 에너지와 밴드갭에 대한 MAPE 감소를 정량화하기 위해 LLaMP 출력과 GPT-3.5 고유 지식을 비교합니다.

실험 결과

연구 질문

  • RQ1LLaMP가 재료 지식 작업에서 고유 지식에 비해 허위 진술을 줄일 수 있는가?
  • RQ2계층화된 ReAct 에이전트가 다중모달 재료 데이터를 (텐서, 결정 구조 등)을 복잡한 질의에 대해 얼마나 잘 검색하고 통합하는가?
  • RQ3LLaMP가 MP 데이터에 대해 합성 절차와 결정 생성 작업을 얼마나 잘 접지하는가?
  • RQ4스칼라 특성(예: 형성 에너지, 밴드갭) 및 텐서 특성에서 RAG가 오류 수정에 어떤 영향을 미치는가?

주요 결과

  • GPT-3.5는 회수 지원 없이 형성 에너지에 대해 큰 MAPE를 보이고 부분적이거나 잘못된 탄성 텐서를 보인다.
  • RAG를 활용한 LLaMP는 여러 특성에서 MP 기준 진실과 출력이 일치하도록 하여 허위 진술을 줄이고 올바른 텐서 및 구조 검색을 가능하게 한다.
  • 형성 에너지에 대한 MAPE가 GPT-3.5 고유 지식에서 LLaMP로 바뀌며 MP-진실값에 근접하게 감소한다.
  • LLaMP는 결정 구조를 생성하거나 편집할 때 올바른 격자 매개변수와 결정 설명을 유지하며, 일반적인 GPT-3.5보다 우수한 성능을 보인다.
  • LLaMP는 MP 데이터에 근거한 합성 절차를 추출해 허위로 추정된 절차와 무관한 전구체를 피한다。

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.