Skip to main content
QUICK REVIEW

[Paper Review] The SEN1-2 Dataset for Deep Learning in SAR-Optical Data Fusion

Michael Schmitt, Lloyd Haydn Hughes|arXiv (Cornell University)|Jul 4, 2018
Advanced Image Fusion Techniques14 references4 citations
TL;DR

This paper introduces the SEN1-2 dataset, a large-scale collection of 282,384 co-registered SAR-optical image patches from Sentinel-1 and Sentinel-2, enabling deep learning research in SAR-optical data fusion. The dataset supports applications like SAR image colorization, image matching, and generating artificial optical images from SAR, with a state-of-the-art 93.98% accuracy in SAR-optical patch matching using a pseudo-siamese CNN.

ABSTRACT

While deep learning techniques have an increasing impact on many technical fields, gathering sufficient amounts of training data is a challenging problem in remote sensing. In particular, this holds for applications involving data from multiple sensors with heterogeneous characteristics. One example for that is the fusion of synthetic aperture radar (SAR) data and optical imagery. With this paper, we publish the SEN1-2 dataset to foster deep learning research in SAR-optical data fusion. SEN1-2 comprises 282,384 pairs of corresponding image patches, collected from across the globe and throughout all meteorological seasons. Besides a detailed description of the dataset, we show exemplary results for several possible applications, such as SAR image colorization, SAR-optical image matching, and creation of artificial optical images from SAR input data. Since SEN1-2 is the first large open dataset of this kind, we believe it will support further developments in the field of deep learning for remote sensing as well as multi-sensor data fusion.

Motivation & Objective

  • To address the scarcity of large-scale, co-registered SAR-optical image pairs for deep learning in remote sensing.
  • To support the development of robust deep learning models for SAR-optical data fusion by providing a globally distributed, seasonally diverse dataset.
  • To enable benchmarking of methods in SAR image colorization, image matching, and synthetic optical image generation.
  • To establish a foundation for future extensions with multi-spectral data and land use/land cover labels.
  • To overcome limitations of existing datasets, such as small size and high patch overlap, by providing a diverse, large-scale, and globally representative dataset.

Proposed method

  • The dataset was constructed using Sentinel-1 GRD VV-polarized SAR images and Sentinel-2 Level-1C top-of-atmosphere reflectance data, both resampled to 20m resolution.
  • Image patches of size 64×64 were extracted from co-located scenes across all global landmasses and all four seasons to ensure temporal and spatial diversity.
  • Precise geolocation was achieved using SRTM or ASTER DEMs and precise orbit data, ensuring sub-pixel alignment between SAR and optical patches.
  • No speckle filtering was applied to preserve raw SAR characteristics, allowing end users to apply their own preprocessing.
  • The dataset was split into deterministic training, validation, and test subsets based on scene and season to ensure unbiased evaluation.
  • Pilot models, including a pseudo-siamese CNN for matching and a pix2pix GAN for image-to-image translation, were trained and evaluated on the dataset to demonstrate its utility.

Experimental results

Research questions

  • RQ1Can a large-scale, globally distributed dataset of co-registered SAR-optical image patches enable more robust and generalizable deep learning models in remote sensing?
  • RQ2How effective is the SEN1-2 dataset for training models in SAR-optical image matching, and what performance can be achieved?
  • RQ3To what extent can generative models trained on SEN1-2 synthesize realistic optical images from SAR inputs?
  • RQ4How does the diversity of global locations and seasonal conditions in the dataset affect model generalization across different environmental conditions?
  • RQ5What are the limitations of using top-of-atmosphere reflectance data instead of bottom-of-atmosphere corrected data for downstream deep learning tasks?

Key findings

  • The SEN1-2 dataset contains 282,384 co-registered SAR-optical image patch pairs collected from diverse global locations and all four seasons, ensuring broad spatial and temporal coverage.
  • A pseudo-siamese convolutional neural network trained on a subset of the dataset achieved a 93.98% accuracy in identifying matching SAR-optical patch pairs, with a low false positive rate of 6.02%.
  • Exemplary results using the pix2pix GAN model demonstrated the feasibility of generating plausible artificial optical images from SAR inputs, with visually coherent results across multiple test cases.
  • The dataset enables unbiased model evaluation due to its deterministic split into training, validation, and test sets based on scene and season, minimizing data leakage.
  • The dataset is the first large-scale open resource of its kind, significantly outperforming prior datasets like SARptical in size and diversity, with 10× more patches and lower overlap.
  • The authors plan to release a version 2 with full multi-spectral Sentinel-2 data and atmospheric correction, enhancing utility for advanced remote sensing applications.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.