[논문 리뷰] PANGAEA: A Global and Inclusive Benchmark for Geospatial Foundation Models
PANGAEA는 Geospatial Foundation Models (GFMs)용으로 전세계적이고 다양한 벤치마크 프로토콜을 도입하고, 여러 GFMs을 감독기준(supervised baselines)과 비교 평가하며, 재현 가능하고 확장 가능한 벤치마킹을 위한 오픈 소스 코드를 제공합니다.
Geospatial Foundation Models (GFMs) have emerged as powerful tools for extracting representations from Earth observation data, but their evaluation remains inconsistent and narrow. Existing works often evaluate on suboptimal downstream datasets and tasks, that are often too easy or too narrow, limiting the usefulness of the evaluations to assess the real-world applicability of GFMs. Additionally, there is a distinct lack of diversity in current evaluation protocols, which fail to account for the multiplicity of image resolutions, sensor types, and temporalities, which further complicates the assessment of GFM performance. In particular, most existing benchmarks are geographically biased towards North America and Europe, questioning the global applicability of GFMs. To overcome these challenges, we introduce PANGAEA, a standardized evaluation protocol that covers a diverse set of datasets, tasks, resolutions, sensor modalities, and temporalities. It establishes a robust and widely applicable benchmark for GFMs. We evaluate the most popular GFMs openly available on this benchmark and analyze their performance across several domains. In particular, we compare these models to supervised baselines (e.g. UNet and vanilla ViT), and assess their effectiveness when faced with limited labeled data. Our findings highlight the limitations of GFMs, under different scenarios, showing that they do not consistently outperform supervised models. PANGAEA is designed to be highly extensible, allowing for the seamless inclusion of new datasets, models, and tasks in future research. By releasing the evaluation code and benchmark, we aim to enable other researchers to replicate our experiments and build upon our work, fostering a more principled evaluation protocol for large pre-trained geospatial models. The code is available at https://github.com/VMarsocci/pangaea-bench.
연구 동기 및 목표
- 좁은 하위 작업 및 지리 편향 데이터셋을 넘어선 GFMs의 견고한 평가를 촉진한다.
- 도시, 농업, 해양, 숲 환경을 포괄하는 다양하고 다도메인 벤치마크를 확립한다.
- 다양한 센서, 해상도 및 시기에 걸쳐 일반화, 데이터 효율성 및 감독 baselines 대비 성능을 평가한다.
- 코드와 모듈식 벤치마킹 프레임워크를 공개하여 재현성과 확장성을 촉진한다.
제안 방법
- 도메인, 양식, 시간성, 지리 정보를 포괄하는 다양한 EO 데이터셋을 선별한다.
- 밀도 예측 태스크(의미론적 분할, 변화 탐지, 회귀)를 포함하되 간단한 패치 수준 분류 및 객체 탐지는 제외한다.
- 자체감독(self-supervised) 및 감독(supervised) 기초 모델을 포함한 다수의 오픈 소스 GFM을 다양한 학습 조건(전부 라벨 vs 제한된 라벨)에서 평가한다.
- 사전 학습 데이터 특성(스펙트럴 풍부성, 공간 해상도)과 다운스트림 태스크/시간 정렬이 GFM 성능에 미치는 영향을 분석한다.
- 새로운 데이터셋, 모델 및 태스크를 추가할 수 있는 확장 가능한 벤치마크 프레임워크를 제공하며 공개 평가 코드를 포함한다.

실험 결과
연구 질문
- RQ1GFMs는 다양한 다운스트림 도메인 및 작업 전반에서 효과적으로 일반화되는가?
- RQ2다양한 센싱 모달리티와 시간 설정에서 GFMs가 지속적으로 감독 기초 대비 우수한 성능을 보이는가?
- RQ3사전 학습 데이터 특성과 라벨 가용성이 GFM 다운스트림 성능에 어떤 영향을 미치는가?
- RQ4태스크와 아키텍처 전반에서 인코더를 미세 조정하는 것과 동결하는 것 사이에 명확한 이점이 있는가?
주요 결과
- GFMs는 일반적으로 다양한 태스크에서 뛰어난 성능을 보이지만 감독 기초 대비 일관되게 우수하지는 않다.
- 사전 학습 데이터 세트에서 스펙트럼 정보가 더 풍부하거나 공간 해상도가 더 높은 경우, 해당 특징이 필요한 태스크의 다운스트림 성능을 높이는 경향이 있다.
- 제한 라벨 환경에서 일부 GFM(CROMA 등)은 일부 기초 모델을 능가할 수 있지만 보편적이지는 않다.
- 미세 조정은 일부 경우 성능을 향상시키지만 인코더를 동결하는 것보다 보편적으로 우수하지는 않다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.