Skip to main content
QUICK REVIEW

[论文解读] A Perceptual Quality Assessment Exploration for AIGC Images

Zicheng Zhang, Chunyi Li|arXiv (Cornell University)|Mar 22, 2023
Image and Video Quality Assessment被引用 4
一句话总结

本文提出了 AGIQA-1K,这是首个针对人工智能生成图像(AGIs)的感知质量评估数据库,包含来自扩散模型的 1,080 幅图像。该研究建立了一个聚焦于技术问题、AI 痕迹、不自然性、差异性以及美学的感知质量评估框架,并对现有图像质量评估(IQA)模型进行了基准测试,揭示了这些模型在 AGIs 上表现有限,原因在于其独特的痕迹和分布偏移。

ABSTRACT

\underline{AI} \underline{G}enerated \underline{C}ontent ( extbf{AIGC}) has gained widespread attention with the increasing efficiency of deep learning in content creation. AIGC, created with the assistance of artificial intelligence technology, includes various forms of content, among which the AI-generated images (AGIs) have brought significant impact to society and have been applied to various fields such as entertainment, education, social media, etc. However, due to hardware limitations and technical proficiency, the quality of AIGC images (AGIs) varies, necessitating refinement and filtering before practical use. Consequently, there is an urgent need for developing objective models to assess the quality of AGIs. Unfortunately, no research has been carried out to investigate the perceptual quality assessment for AGIs specifically. Therefore, in this paper, we first discuss the major evaluation aspects such as technical issues, AI artifacts, unnaturalness, discrepancy, and aesthetics for AGI quality assessment. Then we present the first perceptual AGI quality assessment database, AGIQA-1K, which consists of 1,080 AGIs generated from diffusion models. A well-organized subjective experiment is followed to collect the quality labels of the AGIs. Finally, we conduct a benchmark experiment to evaluate the performance of current image quality assessment (IQA) models.

研究动机与目标

  • 识别并形式化人工智能生成图像(AGIs)独有的主要感知质量维度,包括技术问题、AI 痕迹、不自然性、差异性以及美学。
  • 构建首个感知性 AGI 质量评估数据库 AGIQA-1K,包含 1,080 幅由 stable-diffusion-v2 和 stable-inpainting-v1 扩散模型生成的 AGIs。
  • 通过受控的、有条理的主观实验,收集人类标注的跨定义评估维度的质量标签。
  • 对现有图像质量评估(IQA)模型在 AGIs 上的表现进行基准测试,并评估其在 AGI 特异性质量评估中的适用性。
  • 揭示当前 IQA 模型在应用于 AGIs 时的局限性,并强调开发 AGI 特异性质量评估方法的必要性。

提出的方法

  • 作者基于视觉检查和专家分析,定义了 AGIs 的五个关键感知质量维度:技术问题、AI 痕迹、不自然性、差异性以及美学。
  • 利用两种潜在文本到图像扩散模型——stable-diffusion-v2 和 stable-inpainting-v1——生成 1,080 幅 AGIs,使用涵盖主要物体、次要物体、地点、风格和属性的多样化文本提示。
  • 在受控的实验室环境中开展基于人类受试者的主观实验,使用标准化质量量表对每幅 AGI 在五个质量维度上进行评分。
  • AGIQA-1K 数据库包含 1,080 幅 AGIs 及其对应的主观质量标签,数据按生成模型和图像风格(动漫 vs. 现实主义)划分为子集。
  • 使用标准指标(SRCC、PLCC 和 RMSE)在 AGIQA-1K 上评估 15 种最先进 IQA 模型的性能,涵盖手工设计、手工设计+SVR 以及基于深度学习的方法。
  • 进行统计分析和消融研究,比较模型在完整数据库、模型特定子集以及图像风格类别上的性能表现。
Fig. 1 : Illustration of the generation process of AGIs and NSIs, where NSIs are captured from the natural scenes and AGIs are directly generated from AI models.
Fig. 1 : Illustration of the generation process of AGIs and NSIs, where NSIs are captured from the natural scenes and AGIs are directly generated from AI models.

实验结果

研究问题

  • RQ1哪些主导的感知质量方面使人工智能生成图像(AGIs)与自然场景图像(NSIs)相区别?
  • RQ2AGIs 与 NSIs 在质量相关属性(如模糊、色彩、空间信息)的分布上存在何种差异?
  • RQ3现有 IQA 模型在 AGIs 上的泛化能力如何?其在 AGI 特异性失真上的性能局限性是什么?
  • RQ4扩散模型的选择(如 stable-diffusion-v2 与 stable-inpainting-v1)如何影响质量分布及 IQA 模型的性能?
  • RQ5图像风格(如动漫与现实主义)是否显著影响 IQA 模型在 AGIs 上的表现?

主要发现

  • 基于手工设计的 IQA 模型在 AGIQA-1K 上表现较差,SRCC 值低于 0.05,表明其特征不适用于 AGI 质量表征。
  • 基于深度学习的 IQA 模型表现更高,其中 ResNet50 在完整数据库上取得最佳 SRCC 值 0.6365,但仍未达到令人满意的表现。
  • 所有 IQA 模型在 stable-diffusion-v2 子集上的性能均显著下降,SRCC 从 0.6365 降至 0.4777,可能由于图像多样性与复杂性更高。
  • 在动漫与现实主义风格子集上的表现相似,深度学习模型如 MGQA 分别取得 SRCC 0.6876 和 0.6613,表明风格对模型性能影响有限。
  • AGIs 中质量属性的分布与 NSIs 显著不同,AGIs 展现出更高的模糊度、更多意外痕迹以及更大的不自然性,如归一化概率分布所示。
  • 基准测试表明,当前 IQA 模型无法很好地适应 AGI 特异性失真,凸显了客观 AGI 质量评估中的关键空白。
Fig. 2 : Sample images from the AGIQA-1k database, where the first to sixth rows show AGIs with ( bird, cat, batman, kid, man, woman ) as the main objects respectively.
Fig. 2 : Sample images from the AGIQA-1k database, where the first to sixth rows show AGIs with ( bird, cat, batman, kid, man, woman ) as the main objects respectively.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。