[Paper Review] Foundation Models for Remote Sensing and Earth Observation: A Survey
This survey reviews Remote Sensing Foundation Models (RSFMs), categorizes VFMs, VLMs, LLMs, and other RSFMs, surveys datasets and methods, and outlines challenges and future directions.
Remote Sensing (RS) is a crucial technology for observing, monitoring, and interpreting our planet, with broad applications across geoscience, economics, humanitarian fields, etc. While artificial intelligence (AI), particularly deep learning, has achieved significant advances in RS, unique challenges persist in developing more intelligent RS systems, including the complexity of Earth's environments, diverse sensor modalities, distinctive feature patterns, varying spatial and spectral resolutions, and temporal dynamics. Meanwhile, recent breakthroughs in large Foundation Models (FMs) have expanded AI's potential across many domains due to their exceptional generalizability and zero-shot transfer capabilities. However, their success has largely been confined to natural data like images and video, with degraded performance and even failures for RS data of various non-optical modalities. This has inspired growing interest in developing Remote Sensing Foundation Models (RSFMs) to address the complex demands of Earth Observation (EO) tasks, spanning the surface, atmosphere, and oceans. This survey systematically reviews the emerging field of RSFMs. It begins with an outline of their motivation and background, followed by an introduction of their foundational concepts. It then categorizes and reviews existing RSFM studies including their datasets and technical contributions across Visual Foundation Models (VFMs), Visual-Language Models (VLMs), Large Language Models (LLMs), and beyond. In addition, we benchmark these models against publicly available datasets, discuss existing challenges, and propose future research directions in this rapidly evolving field. A project associated with this survey has been built at https://github.com/xiaoaoran/awesome-RSFMs .
Motivation & Objective
- Motivate the development of RSFMs to handle diverse RS modalities and tasks beyond traditional task-specific models.
- Summarize foundational concepts, architectures, and learning paradigms relevant to RSFMs.
- Catalog and analyze existing RSFM studies, datasets, and technical contributions across VFMs, VLMs, and LLMs.
- Benchmark RSFMs on publicly available RS datasets and discuss limitations and gaps in current research.
- Propose future research directions to advance RSFMs in Earth observation applications.
Proposed method
- categorize RSFMs into VFMs, VLMs, LLMs, and beyond for RS applications.
- review pre-training approaches (supervised, self-supervised) and their alignment with RS data characteristics.
- analyze RS sensor modalities (RGB, MSI, HSI, SAR, LiDAR, DSM, TIR) in RSFMs.
- discuss data sources and pre-training datasets used for RSFMs, including their scale and diversity.
- examine typical RS interpretation tasks (scene classification, semantic segmentation, detection, change detection, VQA, captioning, grounding).
- summarize performance benchmarking and identify challenges and future directions.
Experimental results
Research questions
- RQ1What are the main categories of RSFMs and how do they differ in modalities and downstream capabilities?
- RQ2What datasets and pre-training strategies have been used to develop RSFMs, and how do they impact downstream performance?
- RQ3What are the key challenges in applying foundation models to RS data, and what directions hold promise for future work?
- RQ4How do RSFMs perform across common RS tasks such as scene classification, segmentation, detection, change detection, and VQA?
- RQ5What research directions are recommended to advance RSFMs for diverse EO applications.
Key findings
- RSFMs are categorized into VFMs, VLMs, LLMs, and other RSFMs, each with distinct data modalities and tasks.
- There is a trend toward larger, more diverse RS pre-training datasets, with growing use of self-supervised learning, but scale still lags behind general-domain FMs.
- RS data present domain gaps from natural images, leading to challenges in zero-shot transfer and requiring specialized adaptations.
- SAM and other segmentation-oriented tools are discussed as relevant to RS for promptable segmentation across modalities.
- The survey summarizes multiple RS datasets used for pre-training and benchmarks RSFMs across sensor modalities and tasks.
- Future directions include improving RS-specific architectures, multimodal integration, and scalable RSFM pre-training.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.