Skip to main content
QUICK REVIEW

[논문 리뷰] EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree Representations

Yi-Lun Liao, Brandon C. Wood|arXiv (Cornell University)|2023. 06. 21.
Machine Learning in Materials Science인용 수 81
한 줄 요약

EquiformerV2는 eSCN 컨볼루션과 아키텍처 개선을 이용해 등가 변환기를 더 높은 차수 표현으로 확장하고, OC20에서 속도-정확도 및 데이터 효율이 향상된 최첨단 결과를 달성합니다. 또한 AdsorbML에서 DFT 이완를 줄이고 OC22에서 GemNet-OC 대비 성능을 향상시킵니다.

ABSTRACT

Equivariant Transformers such as Equiformer have demonstrated the efficacy of applying Transformers to the domain of 3D atomistic systems. However, they are limited to small degrees of equivariant representations due to their computational complexity. In this paper, we investigate whether these architectures can scale well to higher degrees. Starting from Equiformer, we first replace $SO(3)$ convolutions with eSCN convolutions to efficiently incorporate higher-degree tensors. Then, to better leverage the power of higher degrees, we propose three architectural improvements -- attention re-normalization, separable $S^2$ activation and separable layer normalization. Putting this all together, we propose EquiformerV2, which outperforms previous state-of-the-art methods on large-scale OC20 dataset by up to $9\%$ on forces, $4\%$ on energies, offers better speed-accuracy trade-offs, and $2 imes$ reduction in DFT calculations needed for computing adsorption energies. Additionally, EquiformerV2 trained on only OC22 dataset outperforms GemNet-OC trained on both OC20 and OC22 datasets, achieving much better data efficiency. Finally, we compare EquiformerV2 with Equiformer on QM9 and OC20 S2EF-2M datasets to better understand the performance gain brought by higher degrees.

연구 동기 및 목표

  • Transformer 기반 아키텍처에서 더 높은 차수의 등가 표현(Lmax)을 가능하게 함으로써 3D 원자모형화 발전.
  • 대규모 OC20/OC22 데이터세트에서 힘과 에너지 예측 정확도 향상.
  • AdsorbML 및 DFT 보강 워크플로우와 같은 실용적 응용을 가능하게 하는 학습 및 추론 효율성 향상.
  • 이전 방법에 비해 데이터 효율성과 OOD 성능 개선을 입증

제안 방법

  • EquiformerV2에서 SO(3) 컨볼루션을 eSCN 컨볼루션으로 대체하여 더 높은 Lmax(최대 6–8) 가능하게 함.
  • 주의 재정규화(attention re-normalization), 분리 가능한 S2 활성화, 분리 가능한 LN 도입 등 아키텍처 개선.
  • 동일하게 등가 그래프 주의(attention) 및 피드포워드 구조를 유지하되 eSCN으로 차수 및 채널 정보를 혼합.
  • 반경 거리 임베딩과 SO(2) 선형 계층을 사용하여 효율적이면서 등가 메시지 전달 구현.
  • OC20 S2EF(All 및 All+MD)와 OC22 데이터셋(AdsorbML 시나리오 포함)에서 평가하고 이전 최첨단과 비교.
Figure 1: Overview of EquiformerV2. We highlight the differences from Equiformer (Liao & Smidt, 2023 ) in red . For (b), (c), and (d), the left figure is the original module in Equiformer, and the right figure is the revised module in EquiformerV2. Input 3D graphs are embedded with atom and edge-deg
Figure 1: Overview of EquiformerV2. We highlight the differences from Equiformer (Liao & Smidt, 2023 ) in red . For (b), (c), and (d), the left figure is the original module in Equiformer, and the right figure is the revised module in EquiformerV2. Input 3D graphs are embedded with atom and edge-deg

실험 결과

연구 질문

  • RQ1효율적인 컨볼루션을 사용하여 3D 원자계 시스템에서 더 높은 차수 표현(Lmax)을 등가 변환기에서 scalable하게 확장할 수 있는가?
  • RQ2아키텍처 개선(주의 정규화, 분리 가능한 S2 활성화, 분리 가능한 LN)이 더 높은 차수 사용 시 기본 Equiformer 대비 측정 가능한 이점을 제공하는가?
  • RQ3EquiformerV2가 OC20/OC22 벤치마크에서 데이터 효율적이고 빠른 학습을 달성하면서 기존 방법을 능가할 수 있는가?
  • RQ4AdsorbML에서 EquiformerV2가 흡착에너지 워크플로우에 미치는 영향 및 DFT 계산 감소 효과는 어떠한가?
  • RQ5QM9 및 OC20 S2EF-2M에서 차수 상승으로 인한 성능 향상 측면에서 EquiformerV2가 Equiformer 및 GemNet-OC와 비교하여 어떤 이점을 보이는가?

주요 결과

  • eSCN 컨볼루션과 더 높은 차수로 구성된 EquiformerV2가 OC20에서 이전 최첨단을 능가하며 힘에서 최대 9%, 에너지에서 최대 4%의 향상을 달성합니다.
  • 주의 재정규화와 분리 가능한 S2 활성화가 학습 안정성과 힘/에너지 정확도를 개선하며, 분리 가능한 LN이 힘 MAE를 추가로 향상시킵니다.
  • OC22에서 단독으로 학습된 EquiformerV2가 OC20 및 OC22에서 학습된 GemNet-OC를 능가하여 강력한 데이터 효율성을 시사합니다.
  • AdsorbML에서 EquiformerV2가 성공률을 높이고 비교 가능한 흡착에너지 정확도에 필요한 DFT 계산을 최대 2배까지 줄이는 효과를 보입니다.
  • λE=4, λF=100인 EquiformerV2(153M 파라미터)가 OC20 S2EF-All+MD에서 최첨단 결과를 달성하며 힘 MAE 향상 및 우수한 속도-정확도 트레이드를 제공합니다.
  • 더 작은 변형(λE=4, λF=100, 31M)도 강력한 성능과 우수한 학습/추론 효율을 보이며 확장 가능한 배치를 시사합니다.
Figure 2: Illustration of different activation functions. $G$ denotes conversion from vectors to point samples on a sphere, $F$ can typically be a SiLU activation or MLPs, and $G^{-1}$ is the inverse of $G$ .
Figure 2: Illustration of different activation functions. $G$ denotes conversion from vectors to point samples on a sphere, $F$ can typically be a SiLU activation or MLPs, and $G^{-1}$ is the inverse of $G$ .

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.