Skip to main content
QUICK REVIEW

[Paper Review] Accurate Building Detection in VHR Remote Sensing Images using Geometric Saliency

Jin Huang, Gui-Song Xia|arXiv (Cornell University)|Jun 4, 2018
Remote-Sensing Image Classification4 citations
TL;DR

This paper proposes a geometric saliency-based method for accurate building detection in very high resolution (VHR) remote sensing images, leveraging mid-level geometric representations through meaningful junctions to compute a novel Geometric Building Index (GBI). The GBI achieves state-of-the-art performance without requiring any training data, outperforming both traditional and learning-based methods in accuracy and generalization across diverse datasets.

ABSTRACT

This paper aims to address the problem of detecting buildings from remote sensing images with very high resolution (VHR). Inspired by the observation that buildings are always more distinguishable in geometries than in texture or spectral, we propose a new geometric building index (GBI) for accurate building detection, which relies on the geometric saliency of building structures. The geometric saliency of buildings is derived from a mid-level geometric representations based on meaningful junctions that can locally describe anisotropic geometrical structures of images. The resulting GBI is measured by integrating the derived geometric saliency of buildings. Experiments on three public datasets demonstrate that the proposed GBI achieves very promising performance, and meanwhile shows impressive generalization capability.

Motivation & Objective

  • Address the challenge of accurate building detection in very high resolution (VHR) remote sensing images where texture and spectral features lose discriminative power.
  • Overcome the limitations of existing methods that rely on texture, spectrum, or morphological features, which degrade in performance at sub-meter resolutions.
  • Develop a method that preserves precise building boundaries and reduces false positives from non-building structures.
  • Achieve strong generalization across diverse datasets without requiring domain-specific training data.
  • Provide a fully unsupervised, training-free alternative to deep learning-based approaches that suffer from overfitting and poor transferability.

Proposed method

  • Represent VHR-RS images using a mid-level geometric representation based on junctions, particularly L-junctions, which capture anisotropic geometric structures of buildings.
  • Detect junctions using the anisotropic-scale junction (ASJ) detector, encoding each junction with location, orientation, length, and significance (ρ) to measure saliency.
  • Compute geometric saliency by combining the local saliency of each junction with its relational context to encode both local and semi-global geometric structure.
  • Define the Geometric Building Index (GBI) as the integrated geometric saliency across the entire image, emphasizing building-like structures.
  • Use prior probability estimates of junctions from one dataset (Spacenet-65) to enable generalization to other datasets without retraining.
  • Leverage the invariance of junctions at building corners regardless of texture or luminance, enabling robust detection under varying visual conditions.

Experimental results

Research questions

  • RQ1Can geometric saliency derived from mid-level junction representations outperform texture- and spectrum-based building detection in VHR-RS images?
  • RQ2To what extent can a training-free, unsupervised method generalize across diverse VHR-RS datasets with varying resolutions and urban structures?
  • RQ3Can geometric saliency effectively preserve accurate building boundaries and reduce false positives from roads, forests, and other non-building features?
  • RQ4How does the performance of a geometric saliency-based method compare to deep learning-based approaches in terms of accuracy and generalization?
  • RQ5Can junction-based geometric features serve as a reliable and robust representation for building detection independent of image texture or spectral content?

Key findings

  • The proposed Geometric Building Index (GBI) achieves the highest mAP (0.46) and F-score (0.59) on the Potsdam dataset among all compared methods in the zero-shot setting.
  • On the Spacenet-65 dataset, GBI achieves mAP of 0.46 and F-score of 0.52, outperforming all other non-learning-based methods and showing strong generalization.
  • Even when trained on the Massachusetts dataset, the HF-FCN model (a learning-based method) achieves very low mAP (0.04) and F-score (0.12) on Spacenet-65 and Potsdam, indicating severe overfitting and poor generalization.
  • GBI maintains high performance across datasets with significant resolution differences (0.5m to 0.05m), demonstrating robustness to resolution changes.
  • Visual results show that GBI detects buildings with clearer, more accurate boundaries and significantly fewer false positives compared to BASI, MBI, and PBI, especially in low-texture or high-clutter regions.
  • The method effectively detects buildings with low-texture roofs and uneven luminance, where texture-based methods like BASI fail, due to the invariant nature of corner junctions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.