[Paper Review] LandCoverNet: A global benchmark land cover classification training dataset
LandCoverNet provides an open-access, globally representative 10m Sentinel-2 based land cover training dataset, with pixel-level labels generated via consensus among three annotators and a helper time-series model.
Regularly updated and accurate land cover maps are essential for monitoring 14 of the 17 Sustainable Development Goals. Multispectral satellite imagery provide high-quality and valuable information at global scale that can be used to develop land cover classification models. However, such a global application requires a geographically diverse training dataset. Here, we present LandCoverNet, a global training dataset for land cover classification based on Sentinel-2 observations at 10m spatial resolution. Land cover class labels are defined based on annual time-series of Sentinel-2, and verified by consensus among three human annotators.
Motivation & Objective
- Define a globally representative land cover taxonomy via expert consensus for Sentinel-2 based classification.
- Create a large-scale, open training dataset with time-series labeled pixels to support global LC mapping.
- Implement a consensus labeling workflow to reduce human error at 10 m resolution.
- Provide a sampling scheme and data specification to enable reproducible benchmarking.
Proposed method
- Define a hierarchical land cover taxonomy through an expert workshop.
- Sample Sentinel-2 tiles (300 total) across continents using MODIS-based class distributions as sampling features.
- Extract 256x256 chips (30 chips per selected tile; ~9000 chips globally; ~589 million pixels).
- Use a time-series based label generation approach with a Random Forest guess label to aid annotators.
- Collect per-pixel labels from three annotators per chip and compute a Bayesian consensus score to determine the final label.
Experimental results
Research questions
- RQ1How to construct a globally representative, open-access land cover training dataset for 10 m Sentinel-2 imagery?
- RQ2What sampling strategy ensures continental diversity and class balance for CHIPS in a global LC benchmark?
- RQ3How can consensus labeling and annotator reliability be used to produce high-quality pixel-level LC labels at 10 m resolution?
Key findings
- LandCoverNet v1.0 covers Africa with 1980 chips (256x256) and excludes permanent snow/ice typical for Africa.
- Consensus scores are generally high; 60% of pixels have a consensus score of 100%.
- The per-pixel consensus score can be used to adjust model training confidence for high- vs. lower-score pixels.
- A time-series based labeling approach using 24 Sentinel-2 scenes per tile and a simple Random Forest guess label facilitated annotator labeling.
- Labels are released under CC BY 4.0 via Radiant MLHub; dataset emphasizes global diversity and open access.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.