Skip to main content
QUICK REVIEW

[논문 리뷰] JARVIS-Leaderboard: A Large Scale Benchmark of Materials Design Methods

Kamal Choudhary, Daniel Wines|arXiv (Cornell University)|2023. 06. 20.
Machine Learning in Materials Science참고 문헌 193인용 수 4
한 줄 요약

JARVIS-Leaderboard는 AI, 전자구조, 힘장, 양자계산 및 실험을 포함한 다양한 재료 설계 방법을 통합한 오픈소스이자 커뮤니티 주도의 벤치마크 플랫폼입니다. 완전한 재료와 결함이 있는 재료를 포함한 여러 데이터 모odal리티에서 작동하며, 메트릭스인 MAE 및 MAD/MAE 비율을 사용하여 5,572개의 재료를 포함한 JARVIS-DFT 3D 데이터셋에서 메서드의 재현성, 투명성 및 표준화된 평가를 가능하게 합니다. 이 플랫폼은 1,200개 이상의 기여를 기록하고 있으며, 다양한 메서드 간의 종합적인 성능 비교를 제공합니다.

ABSTRACT

Lack of rigorous reproducibility and validation are major hurdles for scientific development across many fields. Materials science in particular encompasses a variety of experimental and theoretical approaches that require careful benchmarking. Leaderboard efforts have been developed previously to mitigate these issues. However, a comprehensive comparison and benchmarking on an integrated platform with multiple data modalities with both perfect and defect materials data is still lacking. This work introduces JARVIS-Leaderboard, an open-source and community-driven platform that facilitates benchmarking and enhances reproducibility. The platform allows users to set up benchmarks with custom tasks and enables contributions in the form of dataset, code, and meta-data submissions. We cover the following materials design categories: Artificial Intelligence (AI), Electronic Structure (ES), Force-fields (FF), Quantum Computation (QC) and Experiments (EXP). For AI, we cover several types of input data, including atomic structures, atomistic images, spectra, and text. For ES, we consider multiple ES approaches, software packages, pseudopotentials, materials, and properties, comparing results to experiment. For FF, we compare multiple approaches for material property predictions. For QC, we benchmark Hamiltonian simulations using various quantum algorithms and circuits. Finally, for experiments, we use the inter-laboratory approach to establish benchmarks. There are 1281 contributions to 274 benchmarks using 152 methods with more than 8 million data-points, and the leaderboard is continuously expanding. The JARVIS-Leaderboard is available at the website: https://pages.nist.gov/jarvis_leaderboard

연구 동기 및 목표

  • 재료 과학 분야에서의 재현성과 검증의 부족을 해결하기 위해 통합된 벤치마킹 인fra를 구축합니다.
  • AI, 전자구조, 힘장, 양자계산 및 실험 데이터를 포함한 다양한 재료 설계 방법을 하나의 확장 가능한 플랫폼에 통합합니다.
  • 커뮤니티 기여를 통해 데이터셋, 코드, 메타데이터를 제공함으로써 투명성, 재현성 및 메서드 검증을 향상시킵니다.
  • 표준화된 평가 메트릭스 및 비교 도구(예: Jupyter 노트북)를 제공하여 메서드 간 일관된 성능 평가를 가능하게 합니다.
  • 이상적인 재료뿐 아니라 결함이 있는 재료 데이터를 지원하여 실제 세계의 복잡성을 반영하고 메서드의 강건성을 향상시킵니다.

제안 방법

  • NIST의 JARVIS 인프라 기반으로 중앙집중식 오픈 액세스 플랫폼을 개발하여 AI, 전자구조, 힘장, 양자계산 및 실험 분야의 다섯 영역에서의 벤치마킹 작업을 운영합니다.
  • 평균 절대 오차(MAE) 및 MAE에 대한 평균 절대 편차 비율(MAD/MAE)과 같은 메트릭스를 사용하여 표준화된 평가 프로토콜을 구현하여 메서드 간 비교를 가능하게 합니다.
  • 원자 구조, 원자적 이미지, 스펙트럼, 텍스트, 그리고 다중 실험실 라운드로빈 연구에서 유래한 실험 결과를 포함한 다양한 데이터 모달리티를 통합합니다.
  • AI 벤치마크에서 다양한 입력 유형을 지원하며, Magpie 및 Voronoi-타일링 기반 기술자(특성)를 사용하고 신경망 모델과 비교합니다.
  • Jupyter 및 Google Colab 노트북을 통해 결과의 시각화 및 성능 분석을 가능하게 하여 재현성을 확보합니다.
  • 버전 제어 및 기록 추적 기능을 갖춘 체계적인 기여 프로세스를 통해 데이터셋, 코드, 메타데이터의 기여를 촉진합니다.
Figure 1: Leaderboard snapshot showing an example output for AI based formation energy per atom model on the JARVIS-DFT (dft_3d) dataset. The benchmark has seven contributions so far and they are sorted based on the mean absolute error (MAE) values. Lower MAE values indicate higher accuracy. Links t
Figure 1: Leaderboard snapshot showing an example output for AI based formation energy per atom model on the JARVIS-DFT (dft_3d) dataset. The benchmark has seven contributions so far and they are sorted based on the mean absolute error (MAE) values. Lower MAE values indicate higher accuracy. Links t

실험 결과

연구 질문

  • RQ1다양한 입력 데이터를 기반으로 한 AI 모델들(예: 기술자 기반 vs. 신경망)은 형성 에너지와 같은 재료 성질 예측에서 어떻게 성능을 발휘하는가?
  • RQ2다양한 소프트웨어 패키지, 허위파면(퍼스펙티브) 및 재료에 대해 전자구조 방법의 일관성과 정확도는 실험 데이터와 비교해 볼 때 어떻게 되는가?
  • RQ3다양한 재료에서 부피 모odulus 및 기타 기계적 성질 예측에서 힘장 모델 간의 성능은 어떻게 비교되는가?
  • RQ4다양한 양자 회로 및 알고리즘은 하미르토니안 시뮬레이션을 통해 알루미늄의 전자 밴드 구조를 시뮬레이션할 때 어떤 정도의 정확도를 보이는가?
  • RQ5특히 제올라이트에서의 CO2 흡착 성질과 같은 복잡한 성질에 대해 여러 실험실 간 실험 결과는 얼마나 재현 가능한가?

주요 결과

  • AI 회귀 작업에서 가장 우수한 기술자 기반 모델은 Magpie 및 Voronoi-타일링 기반 특성을 사용한 트리 기반 모델이었으며, 모든 테스트 벤치마크에서 신경망 모델을 능가하는 성능을 보였다.
  • JARVIS-DFT 3D 데이터셋(5,572개 재료 포함)에서 AI 벤치마크는 다양한 입력 모달리티 간의 명확한 메서드 순위를 가능하게 하는 일관된 성능을 보였다.
  • 전자구조 벤치마크에서는 소프트웨어 패키지 및 허위파면에 따라 밴드 갭 예측에 상당한 변동성이 나타났으며, 일부 경우에서 실험 값으로부터 수 eV 이상의 편차가 발생하였다.
  • 양자계산 벤치마크에서는 다양한 양자 회로 및 알고리즘이 알루미늄의 전자 밴드 구조를 시뮬레이션할 때 서로 다른 정확도를 보였으며, 성능은 회로 깊이 및 큐비트 매핑 방식에 따라 달라졌다.
  • ZSM-5 제올라이트에서의 CO2 흡착에 대한 다중 실험실 실험 비교에서 실험실 간에 측정 가능한 변동성이 나타났으며, 이는 표준화된 프로토콜의 필요성을 강조하였다.
  • MAD/MAE 비율은 성능 비교에 있어 강력한 메트릭스로 기능하였으며, 일부 모델이 특히 고차원적 또는 노이즈가 많은 데이터 영역에서 평균 오차 대비 높은 분산을 보이는 것으로 나타났다.
Figure 2: A flow-chart showing the processes involved in uploading a new contribution to the leaderbaord. The jarvis_populate_data.py scripts generate a benchmark dataset. A user can apply their method, train models, or run experiments on that dataset and prepare a csv.zip, a metadata.json file, and
Figure 2: A flow-chart showing the processes involved in uploading a new contribution to the leaderbaord. The jarvis_populate_data.py scripts generate a benchmark dataset. A user can apply their method, train models, or run experiments on that dataset and prepare a csv.zip, a metadata.json file, and

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.