[Paper Review] OpenForest: A data catalogue for machine learning in forest monitoring
OpenForest introduces a dynamic, community-curated catalogue of 86 open access forest datasets spanning ground, aerial, and satellite remote sensing, enabling large-scale machine learning applications in forest monitoring. By centralizing geographically diverse, multi-modal data with standardized metadata, the platform accelerates research in tree species identification, biomass estimation, and foundation model development for global forest sustainability.
Forests play a crucial role in Earth's system processes and provide a suite of social and economic ecosystem services, but are significantly impacted by human activities, leading to a pronounced disruption of the equilibrium within ecosystems. Advancing forest monitoring worldwide offers advantages in mitigating human impacts and enhancing our comprehension of forest composition, alongside the effects of climate change. While statistical modeling has traditionally found applications in forest biology, recent strides in machine learning and computer vision have reached important milestones using remote sensing data, such as tree species identification, tree crown segmentation and forest biomass assessments. For this, the significance of open access data remains essential in enhancing such data-driven algorithms and methodologies. Here, we provide a comprehensive and extensive overview of 86 open access forest datasets across spatial scales, encompassing inventories, ground-based, aerial-based, satellite-based recordings, and country or world maps. These datasets are grouped in OpenForest, a dynamic catalogue open to contributions that strives to reference all available open access forest datasets. Moreover, in the context of these datasets, we aim to inspire research in machine learning applied to forest biology by establishing connections between contemporary topics, perspectives and challenges inherent in both domains. We hope to encourage collaborations among scientists, fostering the sharing and exploration of diverse datasets through the application of machine learning methods for large-scale forest monitoring. OpenForest is available at https://github.com/RolnickLab/OpenForest .
Motivation & Objective
- To address the fragmentation and inaccessibility of open forest datasets across diverse remote sensing platforms and geographic regions.
- To establish a centralized, dynamic, and community-updated repository for open access forest datasets to support machine learning research.
- To inspire interdisciplinary collaboration by linking machine learning challenges with forest monitoring needs and data availability.
- To improve geographical representativeness in forest datasets, particularly for underrepresented regions like Africa and Asia.
- To facilitate the development of multi-modal, foundation models for forest monitoring through curated, harmonized data.
Proposed method
- The OpenForest catalogue aggregates 86 open access forest datasets from inventories, ground-based, aerial, and satellite platforms, categorized by spatial scale, sensor type, and data modality.
- Datasets are curated with standardized metadata, including data provider information, spatial and temporal resolution, and annotation quality, to ensure usability for machine learning.
- The platform is hosted on GitHub and open for community contributions, enabling continuous updates and integration of new datasets like those from the ESA Biomass mission.
- The framework emphasizes multi-modal data alignment—especially LiDAR, SAR, and hyperspectral data—across spatial, temporal, and spectral resolutions to support advanced model training.
- Citizen science data from platforms like OpenAerialMap and GBIF are integrated to enhance global coverage and support self-supervised learning approaches.
- The catalogue supports the development of foundation models by providing diverse, large-scale, and multi-task datasets for pretraining and zero-shot adaptation.
Experimental results
Research questions
- RQ1How can a centralized, community-maintained data catalogue improve access to and usability of open forest datasets for machine learning research?
- RQ2What are the key gaps in geographical and sensor diversity across existing open forest datasets, and how can they be addressed through systematic curation?
- RQ3To what extent can multi-modal, multi-resolution forest data enable the training of robust foundation models for forest monitoring tasks?
- RQ4How can citizen-generated and remote sensing data be harmonized to support self-supervised and transfer learning in forest monitoring?
- RQ5What role does data curation and metadata standardization play in enabling large-scale, cross-scale forest monitoring using machine learning?
Key findings
- The OpenForest catalogue currently hosts 86 open access forest datasets spanning ground-based inventories, aerial surveys, and satellite observations, with comprehensive metadata and provider information.
- The platform enables researchers to identify and access datasets with high spatial and temporal resolution, including multi-modal data such as LiDAR, SAR, and hyperspectral imagery.
- The catalogue highlights significant data gaps in African and Asian regions, emphasizing the need for more geographically representative datasets to reduce bias in machine learning models.
- Integration of upcoming data from the ESA Biomass mission—providing P-band SAR—will significantly enhance global forest tomography and carbon stock estimation capabilities.
- The use of citizen-generated data from GBIF and OpenAerialMap improves global data coverage and supports self-supervised learning, particularly in data-scarce regions.
- The dynamic, community-updated nature of OpenForest fosters collaboration and accelerates the development of foundation models tailored to forest monitoring tasks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.