Skip to main content
QUICK REVIEW

[Paper Review] Image Quality Assessment for Perceptual Image Restoration: A New Dataset, Benchmark and Metric

Jinjin Gu, Haoming Cai|arXiv (Cornell University)|Nov 30, 2020
Advanced Image Processing Techniques50 references31 citations
TL;DR

This paper introduces the PIPAL dataset with GAN-based IR outputs, an Elo-based MOS, and a Space Warping Difference IQA network (SWDN) to benchmark and improve IQA methods for perceptual image restoration. It shows current IQA metrics struggle with GAN-distorted imagery and proposes improvements.

ABSTRACT

Image quality assessment (IQA) is the key factor for the fast development of image restoration (IR) algorithms. The most recent perceptual IR algorithms based on generative adversarial networks (GANs) have brought in significant improvement on visual performance, but also pose great challenges for quantitative evaluation. Notably, we observe an increasing inconsistency between perceptual quality and the evaluation results. We present two questions: Can existing IQA methods objectively evaluate recent IR algorithms? With the focus on beating current benchmarks, are we getting better IR algorithms? To answer the questions and promote the development of IQA methods, we contribute a large-scale IQA dataset, called Perceptual Image Processing ALgorithms (PIPAL) dataset. Especially, this dataset includes the results of GAN-based IR algorithms, which are missing in previous datasets. We collect more than 1.13 million human judgments to assign subjective scores for PIPAL images using the more reliable Elo system. Based on PIPAL, we present new benchmarks for both IQA and SR methods. Our results indicate that existing IQA methods cannot fairly evaluate GAN-based IR algorithms. While using appropriate evaluation methods is important, IQA methods should also be updated along with the development of IR algorithms. At last, we shed light on how to improve the IQA performance on GAN-based distortion. Inspired by the find that the existing IQA methods have an unsatisfactory performance on the GAN-based distortion partially because of their low tolerance to spatial misalignment, we propose to improve the performance of an IQA network on GAN-based distortion by explicitly considering this misalignment. We propose the Space Warping Difference Network, which includes the novel l_2 pooling layers and Space Warping Difference layers. Experiments demonstrate the effectiveness of the proposed method.

Motivation & Objective

  • Motivate evaluation challenges posed by GAN-based perceptual image restoration (IR).
  • Propose a large-scale IQA dataset (PIPAL) including GAN-based distortions.
  • Provide reliable subjective scores via Elo-based MOS and an open rating tool (IQOS).
  • Benchmark existing IQA methods on IR tasks to identify gaps.
  • Propose improvements to IQA for GAN distortions through architecture changes.

Proposed method

  • Create PIPAL with 29k distorted images (including GAN-based outputs) and 116 distortion types across 250 references.
  • Collect subjective scores using the Elo rating system to produce MOS with extensive human judgements (>1.13 million).
  • Develop IQOS, a web-based scoring system integrating Swiss and Elo ratings for scalable annotation.
  • Evaluate a broad suite of FR-IQA and NR-IQA methods (PSNR, SSIM, LPIPS, PieAPP, DISTS, WaDIQaM, NIQE, PI, etc.) against MOS on PIPAL.
  • Analyze GAN-distortion characteristics and identify spatial misalignment as a key weakness of many IQA methods.
  • Propose Space Warping Difference Network (SWDN) with l2-pooling and Space Warping Difference (SWD) layers to improve robustness to misalignment.

Experimental results

Research questions

  • RQ1Can existing IQA methods objectively evaluate GAN-based perceptual IR outputs?
  • RQ2Are current IQA metrics appropriate benchmarks for modern IR algorithms including GAN-based methods?
  • RQ3How do GAN-based distortions differ from traditional distortions in terms of IQA performance?
  • RQ4Can IQA performance be improved by architectural changes that address spatial misalignment in GAN distortions?

Key findings

  • Existing IQA methods struggle to correlate with human judgments on GAN-based distortions in PIPAL.
  • PieAPP, LPIPS and WaDIQaM show relatively better alignment for GAN-based distortions but still underperform.
  • GAN-based distortions reveal a large drop in correlation for traditional FR-IQA metrics (PSNR, SSIM, etc.).
  • Spatial misalignment is a major factor reducing IQA performance on GAN distortions; addressing it improves robustness.
  • SWDN with l2-pooling and SWD layers achieves state-of-the-art performance on GAN-based distortion in IQA tasks.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.