[Paper Review] On the Opportunities and Challenges of Foundation Models for Geospatial Artificial Intelligence
This paper systematically evaluates existing foundation models (FMs) across geospatial domains, showing text-only geospatial tasks benefit from LLMs in zero-/few-shot settings, while multimodal GeoAI tasks still need task-specialized models; it proposes a multimodal FM framework for GeoAI and discusses risks.
Large pre-trained models, also known as foundation models (FMs), are trained in a task-agnostic manner on large-scale data and can be adapted to a wide range of downstream tasks by fine-tuning, few-shot, or even zero-shot learning. Despite their successes in language and vision tasks, we have yet seen an attempt to develop foundation models for geospatial artificial intelligence (GeoAI). In this work, we explore the promises and challenges of developing multimodal foundation models for GeoAI. We first investigate the potential of many existing FMs by testing their performances on seven tasks across multiple geospatial subdomains including Geospatial Semantics, Health Geography, Urban Geography, and Remote Sensing. Our results indicate that on several geospatial tasks that only involve text modality such as toponym recognition, location description recognition, and US state-level/county-level dementia time series forecasting, these task-agnostic LLMs can outperform task-specific fully-supervised models in a zero-shot or few-shot learning setting. However, on other geospatial tasks, especially tasks that involve multiple data modalities (e.g., POI-based urban function classification, street view image-based urban noise intensity classification, and remote sensing image scene classification), existing foundation models still underperform task-specific models. Based on these observations, we propose that one of the major challenges of developing a FM for GeoAI is to address the multimodality nature of geospatial tasks. After discussing the distinct challenges of each geospatial data modality, we suggest the possibility of a multimodal foundation model which can reason over various types of geospatial data through geospatial alignments. We conclude this paper by discussing the unique risks and challenges to develop such a model for GeoAI.
Motivation & Objective
- Assess the performance of existing foundation models on geospatial tasks across multiple subdomains (Geospatial Semantics, Health Geography, Urban Geography, Remote Sensing).
- Identify the advantages and limitations of task-agnostic FMs in GeoAI tasks, especially for multimodal data.
- Discuss challenges and sketch a vision for a multimodal foundation model tailored to GeoAI tasks.
- Highlight risks and considerations in developing and deploying GeoAI foundation models.
Proposed method
- Benchmark several pre-trained foundation models (LLMs, vision, and multimodal) on seven geospatial tasks across four domains.
- Compare FM performance against state-of-the-art fully supervised, task-specific models.
- Use zero-shot and few-shot prompting for text-centric tasks; implement prompts with few-shot examples for topographic/semantic tasks.
- Assess performance on toponym recognition, location description recognition, dementia death time-series forecasting (state and county levels), POI-based urban function classification, street-view image-based noise intensity classification, and RS image scene classification.
- Analyze results to identify modality-specific strengths/weaknesses and the impact of model size and prompting strategy.

Experimental results
Research questions
- RQ1Can existing foundation models match or outperform task-specific models on geospatial semantics tasks in zero-/few-shot settings?
- RQ2Do FMs perform well on health geography, urban geography, and remote sensing tasks, particularly multimodal tasks?
- RQ3What are the main challenges in applying FMs to multimodal GeoAI data, and how might a multimodal GeoFM framework address them?
- RQ4What risks emerge in developing and deploying multimodal GeoAI foundation models?
Key findings
- LLMs can outperform task-specific supervised baselines on text-only geospatial tasks (toponym recognition and location description recognition) in zero-/few-shot settings for several models; GPT-3 and related models show notable gains on some datasets.
- On multimodal GeoAI tasks (e.g., POI-based urban function classification, street-view image-based noise intensity classification, RS image scene classification), existing FMs underperform compared to task-specific models.
- GPT-3, InstructGPT, and some ChatGPT variants show strong performance on dementia-time-series forecasting at state level, sometimes surpassing ARIMA-based baselines in zero-shot settings; GPT-2 family generally underperforms time-series baselines.
- In state-level dementia forecasting, InstructGPT can outperform ARIMA in multiple metrics, while GPT-2 models lag significantly; results at county level show similar trends.
- Overall, multimodal GeoAI remains a key challenge for current FMs, motivating the need for a multimodal GeoAI foundation model with geospatial alignments.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.