[Paper Review] Kuro Siwo: 33 billion $m^2$ under the water. A global multi-temporal satellite dataset for rapid flood mapping
Kuro Siwo is a globally distributed, multi-temporal Synthetic Aperture Radar (SAR) dataset comprising 32 flood events with high-quality manual annotations across 63 billion m² of land, including 12.1 billion m² of flooded or permanent water bodies. It enables rapid flood mapping via supervised and self-supervised deep learning, achieving up to 87% F1-score in water detection and setting strong baselines through the BlackBench benchmark.
Global floods, exacerbated by climate change, pose severe threats to human life, infrastructure, and the environment. Recent catastrophic events in Pakistan and New Zealand underscore the urgent need for precise flood mapping to guide restoration efforts, understand vulnerabilities, and prepare for future occurrences. While Synthetic Aperture Radar (SAR) remote sensing offers day-and-night, all-weather imaging capabilities, its application in deep learning for flood segmentation is limited by the lack of large annotated datasets. To address this, we introduce Kuro Siwo, a manually annotated multi-temporal dataset, spanning 43 flood events globally. Our dataset maps more than 338 billion $m^2$ of land, with 33 billion designated as either flooded areas or permanent water bodies. Kuro Siwo includes a highly processed product optimized for flood mapping based on SAR Ground Range Detected, and a primal SAR Single Look Complex product with minimal preprocessing, designed to promote research on the exploitation of both the phase and amplitude information and to offer maximum flexibility for downstream task preprocessing. To leverage advances in large scale self-supervised pretraining methods for remote sensing data, we augment Kuro Siwo with a large unlabeled set of SAR samples. Finally, we provide an extensive benchmark, namely BlackBench, offering strong baselines for a diverse set of flood events from Europe, America, Africa, Asia and Australia.
Motivation & Objective
- Address the lack of large-scale, high-quality annotated SAR datasets for flood mapping in deep learning.
- Enable rapid, accurate flood mapping under diverse global conditions using SAR data, which operates day and night regardless of weather.
- Provide a benchmark with strong baselines to accelerate research in automated flood detection and response.
- Support both supervised and self-supervised learning by releasing a large unlabeled SAR dataset alongside curated annotations.
- Facilitate improved disaster response and risk assessment by enabling models to generalize across unseen flood events on multiple continents.
Proposed method
- Curate a global, multi-temporal SAR dataset with dual-polarization and elevation data for 32 flood events across Europe, America, Africa, and Australia.
- Conduct meticulous manual annotation of flooded areas and permanent water bodies using high-resolution SAR imagery and cross-verified ground truth from CEMS and other sources.
- Integrate temporal context by including pre-event, during-event, and post-event SAR images to improve model generalization and temporal reasoning.
- Develop BlackBench, a unified benchmark for evaluating flood segmentation models on unseen geographical locations and diverse environmental conditions.
- Enable self-supervised pretraining by releasing a large unlabeled SAR dataset to improve model robustness and transfer learning performance.
- Train and evaluate state-of-the-art models such as FloodViT-Decoder and SNUNet-CD on the Kuro Siwo dataset to establish strong performance baselines.

Experimental results
Research questions
- RQ1Can a globally diverse, multi-temporal SAR dataset with high-quality annotations significantly improve the performance of deep learning models in rapid flood mapping?
- RQ2How does the inclusion of temporal context (pre-, during-, and post-flood) affect model generalization and segmentation accuracy?
- RQ3To what extent can self-supervised pretraining on large unlabeled SAR data improve downstream flood detection performance?
- RQ4Why do models struggle to distinguish between permanent water bodies and flooded areas, especially in dynamic riverine environments?
- RQ5Can models trained on Kuro Siwo generalize effectively to unseen flood events across different continents and environmental conditions?
Key findings
- Kuro Siwo achieves an F1-score of approximately 85% for flooded area detection and 87% for general water detection, demonstrating high annotation quality and model performance.
- Models trained on Kuro Siwo achieve over 82% F1-score in detecting flooded areas and over 85% in binary water detection, even when tested on geographically and climatically diverse, unseen events.
- The inclusion of all available pre-event SAR images significantly improves model performance, highlighting the importance of temporal context in flood segmentation.
- FloodViT-Decoder achieves the highest mean Intersection over Union (IoU) among tested models, while SNUNet-CD shows comparable performance, indicating strong baseline capabilities.
- Significant performance degradation is observed in detecting permanent water bodies, particularly in dynamic river systems, due to challenges in distinguishing them from flooded areas and speckle noise in SAR imagery.
- Qualitative analysis reveals that submerged river islands and small, shifting water patches are difficult to segment accurately, suggesting limitations in current models' ability to resolve fine-scale, dynamic water features under noisy SAR conditions.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.