Skip to main content
QUICK REVIEW

[论文解读] Large-scale Text-to-Image Generation Models for Visual Artists' Creative Works

Hyung-Kwon Ko, Gwanmo Park|arXiv (Cornell University)|Oct 16, 2022
Virtual Reality Applications and Impacts被引用 5
一句话总结

本研究探讨了视觉艺术家如何在创意工作流程中采用大规模文本到图像生成模型(LTGMs),如 DALL-E。通过回顾 72 篇文献并采访来自 35 个领域的 28 名艺术家,研究识别出三个核心角色——自动化、探索与中介,并提出了四项设计准则以提升可用性,包括多模态输入、提示工程支持、模型定制化以及可调节的可变性控制。

ABSTRACT

Large-scale Text-to-image Generation Models (LTGMs) (e.g., DALL-E), self-supervised deep learning models trained on a huge dataset, have demonstrated the capacity for generating high-quality open-domain images from multi-modal input. Although they can even produce anthropomorphized versions of objects and animals, combine irrelevant concepts in reasonable ways, and give variation to any user-provided images, we witnessed such rapid technological advancement left many visual artists disoriented in leveraging LTGMs more actively in their creative works. Our goal in this work is to understand how visual artists would adopt LTGMs to support their creative works. To this end, we conducted an interview study as well as a systematic literature review of 72 system/application papers for a thorough examination. A total of 28 visual artists covering 35 distinct visual art domains acknowledged LTGMs' versatile roles with high usability to support creative works in automating the creation process (i.e., automation), expanding their ideas (i.e., exploration), and facilitating or arbitrating in communication (i.e., mediation). We conclude by providing four design guidelines that future researchers can refer to in making intelligent user interfaces using LTGMs.

研究动机与目标

  • 理解视觉艺术家如何将大规模文本到图像生成模型(LTGMs)整合到其创意流程中。
  • 识别 LTGM 在支持艺术工作流程中所扮演的具体角色,例如自动化、创意拓展与沟通中介。
  • 解决艺术家在采用 LTGM 时面临的挑战,包括提示工程限制、可控性不足以及界面设计缺陷。
  • 为未来提升 LTGM 可用性的智能用户界面提供可操作的设计准则。

提出的方法

  • 在人机交互(HCI)领域对 72 篇关于生成模型的系统与应用论文进行系统性文献回顾,以提取用户、任务与角色的重复主题。
  • 对来自 35 个不同视觉艺术领域的 28 名视觉艺术家进行半结构化访谈,以探索实际采用模式与挑战。
  • 对访谈数据采用双重编码方法——演绎与归纳法,其主题基于文献回顾所得结果。
  • 识别出 LTGM 在艺术工作流程中的重复角色:重复性任务的自动化、新颖创意的探索以及沟通中的中介作用。
  • 基于实证发现提出四项设计准则:可变性控制、领域特定的模型定制化、多模态输入支持以及提示工程辅助。
  • 通过艺术家对可用性、控制力与创意自主权的反馈进行定性分析,验证设计建议。

实验结果

研究问题

  • RQ1哪些视觉艺术家子群体最愿意使用 LTGMs?其采纳受哪些因素影响?
  • RQ2视觉艺术家在工作流程中将 LTGMs 用于哪些类型的创意任务?
  • RQ3LTGM 在支持视觉艺术家方面扮演了哪些关键角色——特别是在自动化、创意生成与沟通方面?
  • RQ4阻碍 LTGM 更深层次融入艺术实践的关键可用性障碍是什么?

主要发现

  • 视觉艺术家广泛认为 LTGMs 是自动化重复性创作任务(如生成初稿或变体)的宝贵工具。
  • LTGMs 被用于通过文本提示生成出人意料或抽象的视觉概念,以拓展创意构思。
  • 艺术家重视 LTGMs 作为沟通中介的角色,尤其是在向可能不理解艺术意图的客户或合作者解释想法时。
  • 主要限制在于提示工程的难度,许多艺术家难以仅通过文字表达复杂或抽象的视觉想法。
  • 艺术家表达了对多模态输入(如草图、语音、手势)的需求,以更准确传达其意图,尤其是在处理生动或抽象图像时。
  • 当前的 LTGM 界面缺乏足够的控制力与定制化能力,导致生成结果与艺术愿景不符时产生挫败感。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。