Skip to main content
QUICK REVIEW

[Paper Review] HYPE: A Benchmark for Human eYe Perceptual Evaluation of Generative Models

Sharon Zhou, Mitchell Gordon|arXiv (Cornell University)|Apr 1, 2019
Visual perception and processing mechanisms59 references72 citations
TL;DR

HYPE establishes two human perceptual benchmarks (time-based and time-free) to reliably measure visual realism of generative models, enabling cost-efficient, reproducible, and separable model comparisons across datasets.

ABSTRACT

Generative models often use human evaluations to measure the perceived quality of their outputs. Automated metrics are noisy indirect proxies, because they rely on heuristics or pretrained embeddings. However, up until now, direct human evaluation strategies have been ad-hoc, neither standardized nor validated. Our work establishes a gold standard human benchmark for generative realism. We construct Human eYe Perceptual Evaluation (HYPE) a human benchmark that is (1) grounded in psychophysics research in perception, (2) reliable across different sets of randomly sampled outputs from a model, (3) able to produce separable model performances, and (4) efficient in cost and time. We introduce two variants: one that measures visual perception under adaptive time constraints to determine the threshold at which a model's outputs appear real (e.g. 250ms), and the other a less expensive variant that measures human error rate on fake and real images sans time constraints. We test HYPE across six state-of-the-art generative adversarial networks and two sampling techniques on conditional and unconditional image generation using four datasets: CelebA, FFHQ, CIFAR-10, and ImageNet. We find that HYPE can track model improvements across training epochs, and we confirm via bootstrap sampling that HYPE rankings are consistent and replicable.

Motivation & Objective

  • Define a gold-standard human benchmark for visual realism of generative models grounded in psychophysics.
  • Provide two evaluation variants (time-based and time-free) that are reliable, separable, and cost-efficient.
  • Demonstrate HYPE's ability to rank models consistently across datasets and sampling methods.
  • Compare HYPE against automated metrics and illustrate its use for tracking progress during training.

Proposed method

  • Two HYPE variants: HYPE_time uses adaptive time constraints to find perceptual thresholds for real vs. fake images.
  • HYPE_infinity (HYPE_\u221e) measures human error rate on 50 real and 50 fake images without time constraints.
  • Images are sampled from models and real datasets to form evaluation sets (K=5000 per model, 5000 real per model).
  • Evaluators pass a qualification task to ensure label quality; $65\%$ accuracy on a 100-image task is required for qualification.
  • Bootstrapping is used to compute 95% confidence intervals and standard deviations for reliability.

Experimental results

Research questions

  • RQ1Can a psychophysics-grounded human benchmark reliably distinguish perceptual realism across GANs and sampling methods?
  • RQ2Do time-based and time-free variants yield consistent rankings and separable model differences?
  • RQ3How does HYPE correlates with or diverge from automated metrics like FID, KID, and precision across datasets and models?
  • RQ4Is HYPE scalable and cost-efficient for large-scale model evaluation and progress tracking during training?
  • RQ5How do results generalize beyond faces to objects and other datasets?

Key findings

  • HYPE_time and HYPE_infinity produce consistent model rankings for unconditional face generation across CelebA-64 and FFHQ-1024.
  • StyleGAN with truncation is the top performer on FFHQ-1024 with a HYPE_time of 363.2 ms and HYPE_infinity of 27.6%.
  • HYPE_infinity provides separable distinctions among models on CelebA-64, even when HYPE_time shows bottoming-out effects for some pairs.
  • HYPE shows strong correlation between HYPE_time and HYPE_infinity (rho = 1.0, p = 0.0), while showing weak or variable correlations with FID and KID across tasks.
  • On ImageNet-5, some classes exhibit separable differences among models, while harder classes show consistently low scores across models, indicating task difficulty impacts perceptual realism.
  • CIFAR-10 results show StyleGAN_trunc beginning to outperform earlier models in human perceptual realism; correlations with automated metrics are moderate or insignificant and vary by model class.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.