Skip to main content
QUICK REVIEW

[Paper Review] UG$^{2+}$ Track 2: A Collective Benchmark Effort for Evaluating and Advancing Image Understanding in Poor Visibility Environments

Ye Yuan, Wenhan Yang|arXiv (Cornell University)|Apr 9, 2019
Image Enhancement TechniquesComputer Science149 references36 citations
TL;DR

This paper introduces UG2+ Track 2, a collective benchmark with three real-world datasets (haze, low light, raindrop) to jointly evaluate and advance object/face detection under poor visibility, and reports baseline results showing significant room for improvement.

ABSTRACT

The UG$^{2+}$ challenge in IEEE CVPR 2019 aims to evoke a comprehensive discussion and exploration about how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. In its second track, we focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. While existing enhancement methods are empirically expected to help the high-level end task, that is observed to not always be the case in practice. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotate objects/faces annotated. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. We expect a large participation from the broad research community to address these challenges together.

Motivation & Objective

  • Motivate robust visual sensing in the wild by evaluating detection under weather and illumination degradations.
  • Provide three real-world datasets with annotations to study haze, under-exposure, and rain effects on detection/recognition.
  • Assess whether low-level vision enhancements help high-level recognition and explore semi-/zero-shot learning settings.
  • Offer baseline results to quantify the gap and drive future methodological innovations.

Proposed method

  • Three sub-challenges capture different poor-visibility conditions: haze (Challenge 2.1), low-light/under-exposure (Challenge 2.2), and raindrop occlusions (Challenge 2.3).
  • New benchmark datasets are provided: RESIDE RTTS for hazy scenes with 4,322 annotated images for training/validation and 2,987 held-out hazy test images across five object categories; DARK FACE with 10,000 under-exposure faces for training/validation and 4,000 testing images; and 1,010 raindrop-image pairs for near-zero-shot training with a 2,495-image held-out test set.
  • Baseline evaluations cascade off-the-shelf restoration/enhancement methods and pretrained detectors to quantify the impact of degradation and restoration on detection performance.
  • Evaluation metric is mean average precision (mAP) with IoU threshold 0.5, with higher IoUs used to break ties when mAPs at 0.5 are equal.

Experimental results

Research questions

  • RQ1How do current object/face detectors perform on real-world hazy, rainy, and low-light images without task-specific adaptation?
  • RQ2Can conventional restoration/enhancement pipelines improve high-level detection/recognition performance under poor visibility?
  • RQ3What is the benefit of leveraging (semi-)supervised or zero-shot training setups for detection in degraded environments?
  • RQ4How do real-world degradations differ from synthetically generated ones in terms of detector performance and evaluation?
  • RQ5What are the limitations and directions for jointly optimizing low-level enhancement with high-level recognition?

Key findings

  • Baseline detectors pretrained on clean data show substantially reduced performance on hazy images, with Mask R-CNN achieving around 41.83 mAP on the hazy validation data and comparable results across other detectors.
  • Dehazing can modestly improve detection (about 1% in mAP on average) when detectors are applied to dehazed images, with Mask R-CNN often yielding the best detection performance among tested detectors.
  • In Sub-challenge 2.1, the held-out test set results are significantly lower than validation, indicating a domain gap and the challenge of real-world hazy scenes (e.g., test mAP figures in the tens rather than high tens).
  • Sub-challenge 2.2 uses the DARK FACE dataset with 10,000 annotated faces for training/validation (43,849 faces) and 4,000 test images (37,711 faces), showing substantial variability in scale, pose, and illumination for face detection under under-exposure.
  • Sub-challenge 2.3 provides 1,010 raindrop pairs for training/validation and a 2,495-image held-out test set; early results indicate zero-shot/unsupervised settings remain challenging for road-object detection under rain-related occlusions.
  • Overall, the paper reports that in Challenge 2.1 and 2.2, winners achieve below 65 MAP, and Challenge 2.3 sees no participants surpassing baseline results, underscoring the high difficulty and room for methodological advances.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.