Skip to main content
QUICK REVIEW

[Paper Review] Organic or Diffused: Can We Distinguish Human Art from AI-generated Images?

Anna Yoo Jeong Ha, Josephine Passananti|arXiv (Cornell University)|Feb 5, 2024
Aesthetic Perception and Analysis7 citations
TL;DR

The paper systematically evaluates automated detectors and human experts in distinguishing human-made art from AI-generated images across multiple styles, models, and adversarial conditions, finding Hive and expert humans offer strongest accuracy but complementary weaknesses.

ABSTRACT

The advent of generative AI images has completely disrupted the art world. Distinguishing AI generated images from human art is a challenging problem whose impact is growing over time. A failure to address this problem allows bad actors to defraud individuals paying a premium for human art and companies whose stated policies forbid AI imagery. It is also critical for content owners to establish copyright, and for model trainers interested in curating training data in order to avoid potential model collapse. There are several different approaches to distinguishing human art from AI images, including classifiers trained by supervised learning, research tools targeting diffusion models, and identification by professional artists using their knowledge of artistic techniques. In this paper, we seek to understand how well these approaches can perform against today's modern generative models in both benign and adversarial settings. We curate real human art across 7 styles, generate matching images from 5 generative models, and apply 8 detectors (5 automated detectors and 3 different human groups including 180 crowdworkers, 4000+ professional artists, and 13 expert artists experienced at detecting AI). Both Hive and expert artists do very well, but make mistakes in different ways (Hive is weaker against adversarial perturbations while Expert artists produce higher false positives). We believe these weaknesses will remain as models continue to evolve, and use our data to demonstrate why a combined team of human and automated detectors provides the best combination of accuracy and robustness.

Motivation & Objective

  • Assess the ability of automated detectors to distinguish human art from AI-generated images across multiple AI models and styles.
  • Evaluate performance of three deployed detectors (Hive, Optic, Illuminarty) and two research detectors (DIRE, DE-FAKE) on unperturbed and perturbed images.
  • Compare performance of three human groups (crowdworkers, professional artists, expert AI-detection artists) in identifying AI-generated art.
  • Analyze how adversarial perturbations affect detector robustness and identify complementary strengths of human and automated detection.

Proposed method

  • Curate a dataset of 280 real human artworks across 7 styles and 350 AI-generated images from 5 diffusion models plus hybrids and upscaled variants.
  • Apply five detectors (Hive, Optic, Illuminarty, DIRE, DE-FAKE) to classify images as human or AI-generated and report probabilistic scores.
  • Conduct three human studies (180 crowdworkers, 4000+ professional artists, 13 expert artists) on image classification with a 5-point Likert-like decision framework.
  • Introduce adversarial perturbations including JPEG compression, Gaussian noise, CLIP-based perturbations, and Glaze-style perturbations to test detector robustness.
  • Evaluate detectors under perturbed conditions and analyze failure modes to propose a combined human-ML detection approach.
Figure 1. Samples from curated test set. Human artwork and subsequent matching images produced by generative AI models.
Figure 1. Samples from curated test set. Human artwork and subsequent matching images produced by generative AI models.

Experimental results

Research questions

  • RQ1Can current automated detectors and human experts reliably distinguish human art from AI-generated images across diverse art styles?
  • RQ2How do adversarial perturbations affect the accuracy of detectors and humans in identifying AI-generated art?
  • RQ3What are the relative strengths and weaknesses of automated versus human detectors, and does a combined approach improve robustness?

Key findings

  • Hive achieves the highest unperturbed accuracy at 98.03% with 0% FPR and 3.17% FNR.
  • Professional and expert artists show high accuracy but introduce more false positives; expert artists detect AI images well but may misclassify human art as AI.
  • Optic and Illuminarty perform worse than Hive, with higher FPRs (24.47% and 67.40%) and varying FNRs.
  • DIRE and DE-FAKE detectors perform poorly on art-specific data, with accuracies around 50% or lower.
  • Adversarial perturbations significantly reduce ML detector performance, especially for feature-space perturbations; CLIP-based perturbations and Glaze perturbations reveal distinct vulnerabilities.
  • A combined team of human and automated detectors yields the best overall accuracy and robustness.
Figure 2. The confidence score produced by automated detectors on images generated by 5 generators. Detecting images generated by Firefly is the hardest.
Figure 2. The confidence score produced by automated detectors on images generated by 5 generators. Detecting images generated by Firefly is the hardest.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.