Skip to main content
QUICK REVIEW

[Paper Review] LandCoverNet: A global benchmark land cover classification training dataset

Hamed Alemohammad, Kevin T. Booth|arXiv (Cornell University)|Dec 5, 2020
Remote-Sensing Image Classification35 citations
TL;DR

LandCoverNet provides an open-access, globally representative 10m Sentinel-2 based land cover training dataset, with pixel-level labels generated via consensus among three annotators and a helper time-series model.

ABSTRACT

Regularly updated and accurate land cover maps are essential for monitoring 14 of the 17 Sustainable Development Goals. Multispectral satellite imagery provide high-quality and valuable information at global scale that can be used to develop land cover classification models. However, such a global application requires a geographically diverse training dataset. Here, we present LandCoverNet, a global training dataset for land cover classification based on Sentinel-2 observations at 10m spatial resolution. Land cover class labels are defined based on annual time-series of Sentinel-2, and verified by consensus among three human annotators.

Motivation & Objective

  • Define a globally representative land cover taxonomy via expert consensus for Sentinel-2 based classification.
  • Create a large-scale, open training dataset with time-series labeled pixels to support global LC mapping.
  • Implement a consensus labeling workflow to reduce human error at 10 m resolution.
  • Provide a sampling scheme and data specification to enable reproducible benchmarking.

Proposed method

  • Define a hierarchical land cover taxonomy through an expert workshop.
  • Sample Sentinel-2 tiles (300 total) across continents using MODIS-based class distributions as sampling features.
  • Extract 256x256 chips (30 chips per selected tile; ~9000 chips globally; ~589 million pixels).
  • Use a time-series based label generation approach with a Random Forest guess label to aid annotators.
  • Collect per-pixel labels from three annotators per chip and compute a Bayesian consensus score to determine the final label.

Experimental results

Research questions

  • RQ1How to construct a globally representative, open-access land cover training dataset for 10 m Sentinel-2 imagery?
  • RQ2What sampling strategy ensures continental diversity and class balance for CHIPS in a global LC benchmark?
  • RQ3How can consensus labeling and annotator reliability be used to produce high-quality pixel-level LC labels at 10 m resolution?

Key findings

  • LandCoverNet v1.0 covers Africa with 1980 chips (256x256) and excludes permanent snow/ice typical for Africa.
  • Consensus scores are generally high; 60% of pixels have a consensus score of 100%.
  • The per-pixel consensus score can be used to adjust model training confidence for high- vs. lower-score pixels.
  • A time-series based labeling approach using 24 Sentinel-2 scenes per tile and a simple Random Forest guess label facilitated annotator labeling.
  • Labels are released under CC BY 4.0 via Radiant MLHub; dataset emphasizes global diversity and open access.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.