Skip to main content
QUICK REVIEW

[Paper Review] A Perceptual Quality Assessment Exploration for AIGC Images

Zicheng Zhang, Chunyi Li|arXiv (Cornell University)|Mar 22, 2023
Image and Video Quality Assessment4 citations
TL;DR

This paper introduces AGIQA-1K, the first perceptual quality assessment database for AI-generated images (AGIs), comprising 1,080 images from diffusion models. It establishes a quality assessment framework focusing on technical issues, AI artifacts, unnaturalness, discrepancy, and aesthetics, and benchmarks existing IQA models, revealing their limited performance on AGIs due to unique artifacts and distribution shifts.

ABSTRACT

\underline{AI} \underline{G}enerated \underline{C}ontent ( extbf{AIGC}) has gained widespread attention with the increasing efficiency of deep learning in content creation. AIGC, created with the assistance of artificial intelligence technology, includes various forms of content, among which the AI-generated images (AGIs) have brought significant impact to society and have been applied to various fields such as entertainment, education, social media, etc. However, due to hardware limitations and technical proficiency, the quality of AIGC images (AGIs) varies, necessitating refinement and filtering before practical use. Consequently, there is an urgent need for developing objective models to assess the quality of AGIs. Unfortunately, no research has been carried out to investigate the perceptual quality assessment for AGIs specifically. Therefore, in this paper, we first discuss the major evaluation aspects such as technical issues, AI artifacts, unnaturalness, discrepancy, and aesthetics for AGI quality assessment. Then we present the first perceptual AGI quality assessment database, AGIQA-1K, which consists of 1,080 AGIs generated from diffusion models. A well-organized subjective experiment is followed to collect the quality labels of the AGIs. Finally, we conduct a benchmark experiment to evaluate the performance of current image quality assessment (IQA) models.

Motivation & Objective

  • To identify and formalize the major perceptual quality aspects unique to AI-generated images (AGIs), including technical issues, AI artifacts, unnaturalness, discrepancy, and aesthetics.
  • To construct the first perceptual AGI quality assessment database, AGIQA-1K, with 1,080 AGIs generated from stable-diffusion-v2 and stable-inpainting-v1 diffusion models.
  • To conduct a controlled, well-organized subjective experiment to collect human-annotated quality labels across the defined assessment dimensions.
  • To benchmark the performance of existing image quality assessment (IQA) models on AGIs and evaluate their suitability for AGI-specific quality evaluation.
  • To reveal the limitations of current IQA models when applied to AGIs and highlight the need for AGI-specific quality assessment methods.

Proposed method

  • The authors define five key perceptual quality dimensions for AGIs: technical issues, AI artifacts, unnaturalness, discrepancy, and aesthetics, based on visual inspection and expert analysis.
  • They generate 1,080 AGIs using two latent text-to-image diffusion models—stable-diffusion-v2 and stable-inpainting-v1—using diverse text prompts covering main objects, second objects, places, styles, and attributes.
  • A controlled laboratory-based subjective experiment is conducted with human subjects who rate each AGI on the five quality dimensions using a standardized quality scale.
  • The AGIQA-1K database is constructed with 1,080 AGIs and corresponding subjective quality labels, with data split into subsets based on generation model and image style (anime vs. realistic).
  • The performance of 15 state-of-the-art IQA models—spanning handcrafted, handcrafted+SVR, and deep learning-based approaches—is evaluated on AGIQA-1K using standard metrics: SRCC, PLCC, and RMSE.
  • Statistical analysis and ablation studies are performed to compare model performance across the full database, model-specific subsets, and image style categories.
Fig. 1 : Illustration of the generation process of AGIs and NSIs, where NSIs are captured from the natural scenes and AGIs are directly generated from AI models.
Fig. 1 : Illustration of the generation process of AGIs and NSIs, where NSIs are captured from the natural scenes and AGIs are directly generated from AI models.

Experimental results

Research questions

  • RQ1What are the dominant perceptual quality aspects that distinguish AI-generated images (AGIs) from natural scene images (NSIs)?
  • RQ2How do the distributions of quality-related attributes (e.g., blur, color, spatial information) differ between NSIs and AGIs?
  • RQ3To what extent do existing IQA models generalize to AGIs, and what are their performance limitations on AGI-specific distortions?
  • RQ4How does the choice of diffusion model (e.g., stable-diffusion-v2 vs. stable-inpainting-v1) affect the quality distribution and IQA model performance?
  • RQ5Does image style (e.g., anime vs. realistic) significantly influence the performance of IQA models on AGIs?

Key findings

  • Handcrafted-based IQA models perform poorly on AGIQA-1K, with SRCC values below 0.05, indicating their features are ill-suited for AGI quality representation.
  • Deep learning-based IQA models achieve higher performance, with ResNet50 achieving the best SRCC of 0.6365 on the full database, but still fall short of satisfactory performance.
  • All IQA models show significant performance drops on the stable-diffusion-v2 subset, with SRCC dropping from 0.6365 to 0.4777, likely due to higher image diversity and complexity.
  • Performance on the anime and realistic style subsets is similar, with deep learning models like MGQA achieving SRCC of 0.6876 and 0.6613 respectively, indicating style has limited impact on model performance.
  • The distribution of quality attributes in AGIs differs significantly from NSIs, with AGIs showing higher blur, more unexpected artifacts, and greater unnaturalness, as shown in normalized probability distributions.
  • The benchmark reveals that current IQA models are not well-adapted to AGI-specific distortions, highlighting a critical gap in objective AGI quality assessment.
Fig. 2 : Sample images from the AGIQA-1k database, where the first to sixth rows show AGIs with ( bird, cat, batman, kid, man, woman ) as the main objects respectively.
Fig. 2 : Sample images from the AGIQA-1k database, where the first to sixth rows show AGIs with ( bird, cat, batman, kid, man, woman ) as the main objects respectively.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.