[논문 리뷰] ClimaX: A foundation model for weather and climate
ClimaX는 이질적인 CMIP6 기후 데이터에 대해 사전 학습된 유연한 트랜스포머 기반의 기본 모델로, 다양한 기상 및 기후 작업에 대한 미세 조정을 가능하게 하며 탁월한 일반화 및 경쟁력 있는 벤치마크를 제공합니다.
Most state-of-the-art approaches for weather and climate modeling are based on physics-informed numerical models of the atmosphere. These approaches aim to model the non-linear dynamics and complex interactions between multiple variables, which are challenging to approximate. Additionally, many such numerical models are computationally intensive, especially when modeling the atmospheric phenomenon at a fine-grained spatial and temporal resolution. Recent data-driven approaches based on machine learning instead aim to directly solve a downstream forecasting or projection task by learning a data-driven functional mapping using deep neural networks. However, these networks are trained using curated and homogeneous climate datasets for specific spatiotemporal tasks, and thus lack the generality of numerical models. We develop and demonstrate ClimaX, a flexible and generalizable deep learning model for weather and climate science that can be trained using heterogeneous datasets spanning different variables, spatio-temporal coverage, and physical groundings. ClimaX extends the Transformer architecture with novel encoding and aggregation blocks that allow effective use of available compute while maintaining general utility. ClimaX is pre-trained with a self-supervised learning objective on climate datasets derived from CMIP6. The pre-trained ClimaX can then be fine-tuned to address a breadth of climate and weather tasks, including those that involve atmospheric variables and spatio-temporal scales unseen during pretraining. Compared to existing data-driven baselines, we show that this generality in ClimaX results in superior performance on benchmarks for weather forecasting and climate projections, even when pretrained at lower resolutions and compute budgets. The source code is available at https://github.com/microsoft/ClimaX.
연구 동기 및 목표
- 날씨와 기후를 위한 일반적이고 데이터 주도적인 기본 모델의 생성을 촉진하여 작업별 한계를 극복한다.
- 일반성 강화를 위해 이질적이고 물리정보가 반영된 기후 데이터셋(CMIP6)을 사전 학습에 활용한다.
- 가변 모달리티와 불규칙한 시공간 커버리지를 가능하게 하는 아키텍처 혁신을 개발한다.
- 예보, 기후 전망, 다운스케일링 작업 전반에 걸친 미세 조정의 다재다능성을 입증한다.
- 데이터, 모델 크기, 해상도에 따른 스케일링 동학을 평가하여 향후 연구를 안내한다.
제안 방법
- Vision Transformer (ViT)을 두 가지 핵심 모듈로 확장한다: variable tokenization(각 입력 변수별로 토큰화)과 variable aggregation(각 공간 위치에서 크로스 어텐션으로 변수를 융합)이다.
- 임의의 시점까지의 시간 범위를 가이드하는 위치 및 리드타임 임베딩을 추가한다.
- 리드타임 6h–168h에 걸쳐 현재 상태에서 미래 상태를 예측하도록 무작위 예측 목표로 CMIP6 데이터에서 ClimaX를 사전 학습한다.
- Global 그리딩의 구면 면적 차이를 반영하기 위해 위도 가중 평균 제곱 오차를 사용한다.
- 사전 학습 중에 본 변수 및 보지 못한 변수를 포함한 다운스트림 작업에 대해 모듈형 미세조정 전략으로 ClimaX를 미세조정한다.
- 예보, 기후 전망, 다운스케일링 벤치마크에서 평가하고 신경망 벤치마크 및 NWP 벤치마크와 비교한다.
실험 결과
연구 질문
- RQ1단일 사전 학습된 ClimaX 모델이 이질적인 데이터로 다수의 기상 및 기후 작업에 적응할 수 있는가?
- RQ2전세계/지역 예보, 부분계절에서 계절 예측, 기후 전망 및 다운스케일링에서 전문 기반 대비 ClimaX가 어떤 성능을 보이는가?
- RQ3데이터셋 크기, 모델 용량, 해상도가 ClimaX의 성능(스케일링 법칙)에 미치는 영향은 무엇인가?
- RQ4미세조정 전략 중 어떤 것이 ClimaX를 보지 못한 변수와 다양한 시공간 해상도에 가장 잘 이전시키는가?
주요 결과
- 단일 사전 학습된 ClimaX는 서로 다른 리드 타임, 해상도 및 지역에 걸친 광범위한 작업에 대해 미세조정될 수 있다.
- ClimaX는 ClimateBench에서 최첨단 성능을 달성하고 WeatherBench의 운영 IFS와도 상당히 경쟁력이 있으며, 중간 정도의 계산으로도 달성된다.
- 스케일링 분석은 더 많은 사전 학습 데이터, 더 큰 모델, 그리고 더 높은 해상도 입력에서 성능 향상을 보여준다.
- 발췌 연구는 보지 못한 변수 처리와 공유 구성요소가 있는 여러 작업을 포함한 효과적인 미세조정 전략을 뒷받침한다.
- 가변 토큰화와 집계 덕분에 불규칙한 데이터 커버리지와 변수 세트에 대한 견고함을 유지한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.