[논문 리뷰] A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools
본 연구는 기초 모델, LLM 에이전트, 데이터 세트 및 도구가 6개 작업 영역에 걸친 AI 기반 재료과학을 가능하게 하는 방법을 검토하고, 단일 모달, 다중 모달 및 에이전트 기반 모델과 향후 과제를 강조합니다.
Foundation models (FMs) are catalyzing a transformative shift in materials science (MatSci) by enabling scalable, general-purpose, and multimodal AI systems for scientific discovery. Unlike traditional machine learning models, which are typically narrow in scope and require task-specific engineering, FMs offer cross-domain generalization and exhibit emergent capabilities. Their versatility is especially well-suited to materials science, where research challenges span diverse data types and scales. This survey provides a comprehensive overview of foundation models, agentic systems, datasets, and computational tools supporting this growing field. We introduce a task-driven taxonomy encompassing six broad application areas: data extraction, interpretation and Q\&A; atomistic simulation; property prediction; materials structure, design and discovery; process planning, discovery, and optimization; and multiscale modeling. We discuss recent advances in both unimodal and multimodal FMs, as well as emerging large language model (LLM) agents. Furthermore, we review standardized datasets, open-source tools, and autonomous experimental platforms that collectively fuel the development and integration of FMs into research workflows. We assess the early successes of foundation models and identify persistent limitations, including challenges in generalizability, interpretability, data imbalance, safety concerns, and limited multimodal fusion. Finally, we articulate future research directions centered on scalable pretraining, continual learning, data governance, and trustworthiness.
연구 동기 및 목표
- 크로스 도메인의 데이터 풍부한 재료 발견과 설계를 가능하게 하기 위해 기초 모델의 활용을 고무한다.
- 작업, 아키텍처 및 사전 학습 전략에 따라 단일 모달, 다중 모달 및 에이전트 기반 기초 모델을 분류한다.
- 재료과학에서 AI 워크플로우를 지원하는 데이터세트, 도구 및 자율 플랫폼을 요약한다.
제안 방법
- 데이터 추출, 해석 및 Q&A; 원자 수준 시뮬레이션; 특성 예측; 재료 구조, 설계 및 발견; 공정 계획, 발견 및 최적화; 다스케일 모델링의 여섯 가지 응용 영역을 맥락에 맞춘 분류 체계를 제공합니다.
- 단일 모달, 다중 모달 및 에이전트 기반 기초 모델과 그들의 교차 도메인 역량을 검토합니다.
- 대표 모델, 데이터세트 및 도구 체인을 요약하고, 성공 사례, 한계 및 향후 방향을 논의합니다.
실험 결과
연구 질문
- RQ1재료과학에서 AI를 이끄는 주요 기초 모델 아키텍처와 사전 학습 전략은 무엇인가?
- RQ2단일 모달, 다중 모달 및 에이전트 기반 모델이 핵심 MatSci 작업 및 재료 계급에서 어떻게 성능을 발휘하는가?
- RQ3재료 발견 및 설계에서 확장 가능하고 자율적인 AI 워크플로우를 지원하는 데이터세트와 도구는 무엇인가?
- RQ4실제 적용을 저해하는 지속적인 한계와 안전성 우려는 무엇인가?
- RQ5MatSci AI에서 확장 가능한 사전 학습, 지속 학습, 데이터 거버넌스 및 신뢰성 향상을 가능하게 할 미래 방향은 무엇인가?
주요 결과
- GNoME은 그래프 신경망을 활성 학습 기반의 DFT 검증과 결합하여 220만 건이 넘는 새로운 안정한 재료를 발견했다.
- MatterSim은 17 million DFT-labeled structures 고 학습되었으며 모든 원소에 대한 보편 시뮬레이션과 다양한 온도·압력을 지원한다.
- MACE-MP-0 구는 주기적 시스템에 대해 최첨단 정확도를 달성하면서 등가 불변성 유도 편향을 보존한다.
- 다중모달 및 교차 도메인 모델(nach0, MultiMat, MatterChat 등)은 구조, 텍스트 및 분광 데이터에 대한 추론을 가능하게 한다.
- LLM 에이전트(HoneyComb, MatAgent, ChatMOF, MatPilot 등)는 문헌 검토, 설계, 합성 계획 및 실험 워크플로우 전반에 걸쳐 자율적 또는 반자율적 작업을 가능하게 한다.
- 도구 키트(Open MatSci ML Toolkit, FORGE 등)와 자율 플랫폼(A-Lab 등)의 성장하는 생태계가 통합 AI 기반 재료 워크플로우를 지원한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.