[Paper Review] Pinning "Reflection" on the Agenda: Investigating Reflection in Human-LLM Co-Creation for Creative Coding
This study investigates how human-AI collaboration via large language models (LLMs) influences artists' reflection during creative coding, comparing full-program and subtask-based prompting. It finds that reflection type correlates with performance, satisfaction, and subjective experience, revealing that LLM-assisted collaboration enhances creativity when reflection is intentionally supported, offering design implications for future AI-human creative tools.
Large language models (LLMs) are increasingly integrated into creative coding, yet how users reflect, and how different co-creation conditions influence reflective behavior, remains underexplored. This study investigates situated, moment-to-moment reflection in creative coding under two prompting strategies: the entire task invocation (T1) and decomposed subtask invocation (T2), to examine their effects on reflective behavior. Our mixed-method results reveal three distinct reflection types and show that T2 encourages more frequent, strategic, and generative reflection, fostering diagnostic reasoning and goal redefinition. These findings offer insights into how LLM-based tools foster deeper creative engagement through structured, behaviorally grounded reflection support.
Motivation & Objective
- To investigate how different LLM collaboration methods affect artists' reflection during creative coding.
- To examine the correlation between reflection types and user performance, satisfaction, and subjective experience.
- To identify design implications for AI-assisted creative programming tools from the artist’s perspective.
- To address the gap in understanding reflection as a core creative mechanism in human-LLM collaboration.
Proposed method
- Conducted a mixed-methods study with 22 artists of varying programming backgrounds.
- Implemented two collaboration conditions: full-program prompting (T1) and subtask-based prompting (T2) using a custom LLM (Llama we built).
- Collected quantitative data on task performance and user satisfaction across both conditions.
- Conducted semi-structured interviews to analyze qualitative aspects of reflection, user experience, and creative process.
- Categorized reflection types based on observed user behaviors and linguistic patterns in dialogue.
- Used thematic analysis to identify patterns in reflection, design process, and multimodal integration needs.
Experimental results
Research questions
- RQ1RQ1: What are the reflection types and patterns of artists in different ways of collaborating on LLM-assisted programming?
- RQ2RQ2: What kind of collaboration and reflection enables artists to achieve more efficient task performance and better user satisfaction?
- RQ3RQ3: What are artists’ subjective experiences of LLM-assisted programming in different collaboration approaches?
Key findings
- Reflection types varied significantly between full-program and subtask-based collaboration, with subtask prompting eliciting more deliberate and iterative reflection.
- Artists reported higher satisfaction and better task performance in subtask-based collaboration, though reflection sometimes reduced efficiency.
- A strong correlation was found between reflective behaviors and user satisfaction, indicating reflection enhances perceived quality of collaboration.
- Participants expressed a need for multimodal integration, particularly visual cues and semantic context, to improve LLM understanding of creative intent.
- The study revealed a mismatch between performance efficiency and user experience, where reflective engagement improved satisfaction despite slower progress.
- Design implications were derived, including the need for LLMs to support scope, scale, and access flexibility, and to integrate visual and semantic information for better creative alignment.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.