[Paper Review] The Creativity of Text-to-Image Generation
This paper argues that the traditional product-centered view of creativity fails to capture the full scope of human creativity in text-to-image generation, proposing instead a process- and context-centered approach using Rhodes' '4 P' model (product, person, process, press). It highlights prompt engineering and online communities as central to creative practice, demonstrating that creativity emerges from human-AI interaction and community-driven learning rather than just the final image output.
Text-guided synthesis of images has made a giant leap towards becoming a mainstream phenomenon. With text-to-image generation systems, anybody can create digital images and artworks. This provokes the question of whether text-to-image generation is creative. This paper expounds on the nature of human creativity involved in text-to-image art (so-called "AI art") with a specific focus on the practice of prompt engineering. The paper argues that the current product-centered view of creativity falls short in the context of text-to-image generation. A case exemplifying this shortcoming is provided and the importance of online communities for the creative ecosystem of text-to-image art is highlighted. The paper provides a high-level summary of this online ecosystem drawing on Rhodes' conceptual four P model of creativity. Challenges for evaluating the creativity of text-to-image generation and opportunities for research on text-to-image generation in the field of Human-Computer Interaction (HCI) are discussed.
Motivation & Objective
- To challenge the product-centered definition of creativity in text-to-image generation, which reduces creativity to the originality and effectiveness of the final image.
- To demonstrate that the creative process in text-to-image art is not captured by evaluating only the output, but requires attention to the human-AI interaction and community context.
- To highlight the role of online communities in shaping and supporting novel creative practices such as prompt engineering and curation.
- To advocate for a shift in evaluation frameworks toward including process, person, and press (environment) in assessing creativity in AI-generated art.
- To identify opportunities for Human-Computer Interaction (HCI) research in designing co-creative systems that support prompt engineering and community collaboration.
Proposed method
- Adopted Rhodes’ '4 P' model of creativity—product, person, process, and press—as a conceptual framework to analyze human creativity in text-to-image generation.
- Analyzed case studies of text-to-image art creation, focusing on prompt engineering as a core creative practice that involves iterative refinement and experimentation.
- Examined the role of online communities (e.g., Midjourney, Colab, social media) as critical environments (press) that enable knowledge sharing, feedback loops, and collective learning.
- Explored image-level and portfolio-level curation as creative acts that reflect the creator’s intent, aesthetic judgment, and evolving style.
- Drew parallels with natural language prompting in large language models (e.g., GPT-3) to inform design principles for prompt engineering in text-to-image systems.
- Proposed that future HCI research should focus on designing user interfaces and tools that support co-creative workflows, such as MindsEye or community-driven prompt templates.
Experimental results
Research questions
- RQ1How does the product-centered view of creativity fail to capture the full extent of human creativity in text-to-image generation?
- RQ2What role do online communities play in shaping the creative process and practices such as prompt engineering in text-to-image art?
- RQ3In what ways does the interaction between humans and AI systems constitute a form of co-creation that goes beyond the final image output?
- RQ4How can the '4 P' model (product, person, process, press) be applied to better understand and evaluate creativity in AI-generated art?
- RQ5What opportunities exist for HCI research in designing systems that support and enhance prompt engineering and collaborative creativity in text-to-image generation?
Key findings
- The product-centered definition of creativity—requiring originality and effectiveness of the final image—fails to account for the creative labor involved in prompt engineering and iterative refinement.
- Creativity in text-to-image generation is not solely in the image output but is deeply embedded in the process of crafting and refining prompts through trial, error, and community feedback.
- Online communities function as essential 'press' elements in the 4 P model, providing shared resources, templates, and collaborative learning environments that shape creative practice.
- Image-level and portfolio-level curation are significant creative acts that reflect the artist’s aesthetic judgment and evolving identity, extending creativity beyond prompt creation.
- Prompt engineering has emerged as a distinct creative craft practiced within online communities, analogous to traditional artistic techniques but adapted to natural language interaction with AI.
- The rise of text-to-image generation signals a potential societal shift in creative production, comparable to the historical impact of photography, with implications for how we define authorship, skill, and creativity in the digital age.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.