[Paper Review] From Prompts to Worlds: How Users Iterate, Explore, and Make Sense of AI-Generated 3D Environments
This study empirically investigates how users interact with a commercial text-to-3D platform, revealing a language-to-space gap, episodic presence, and iteration barriers that shape sensemaking in AI-created 3D environments.
Text-to-3D generative AI systems create navigable environments from natural language prompts, but unlike text-to-image generation, evaluation requires embodied exploration of spatial coherence, scale, and navigability. We present the first empirical study of a commercial text-to-3D platform, combining think-aloud protocols, behavioral observation, and validated measures of usability, presence, and engagement. We report three findings. First, asymmetric expressibility: users readily convey semantic intent (themes, atmosphere) but struggle to specify spatial structure (layout, scale), reflecting a language-to-space limitation rather than a skill deficit. Second, episodic presence: immersion arises when expectations align with outputs but does not accumulate into sustained place illusion. Third, structural iteration breakdowns: refinement fails due to interaction barriers - poor discoverability, opaque feedback, and high temporal costs - rather than user limitations. Together, these dynamics form a reinforcing cycle in which spatial mismatches persist, producing episodic presence and ongoing sensemaking. We reframe text-to-3D interaction as negotiated meaning-making rather than linear prompting, and argue that effective systems require hybrid input modalities, transparent feedback, and low-cost iteration.
Motivation & Objective
- Motivate understanding of how users translate natural language prompts into navigable 3D spaces.
- Examine how users iterate, explore, and make sense of AI-generated 3D environments through embodied tasks.
- Identify cognitive and interaction barriers that affect usability, presence, and engagement in text-to-3D systems.
- Propose design implications for hybrid input modalities, transparent feedback, and low-cost iteration to improve user experience.
Proposed method
- Combine think-aloud protocols with behavioral observation during interaction with a commercial text-to-3D platform.
- Use validated measures of usability, presence, and engagement to assess user experience.
- Analyze how semantic intent is expressed versus spatial structure specification to identify language-to-space limitations.
- Characterize episodes of presence and their relation to expectation alignment with outputs.
- Identify points of refinement breakdown and their underlying causes, such as discoverability and feedback opacity.
Experimental results
Research questions
- RQ1How do users express semantic intent and spatial structure when using text-to-3D prompts?
- RQ2What are the patterns of presence and immersion in AI-generated 3D environments, and how do they relate to alignment between expectations and outputs?
- RQ3What interaction barriers impede systematic refinement and iteration in text-to-3D tools?
- RQ4What design changes could mitigate language-to-space gaps and support lower-cost, more transparent iteration?
- RQ5How should text-to-3D systems be framed in terms of meaning-making rather than linear prompting?
Key findings
- Users readily convey semantic themes and atmosphere but struggle with specifying spatial layout and scale.
- Immersion (episodic presence) arises when outputs align with expectations but does not accumulate into sustained place illusion.
- Refinement breakdowns stem from interaction barriers such as poor discoverability, opaque feedback, and high temporal costs.
- A reinforcing cycle emerges where spatial mismatches persist, leading to episodic presence and ongoing sensemaking.
- The study argues for viewing text-to-3D interaction as negotiated meaning-making and suggests hybrid inputs, transparent feedback, and low-cost iteration.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.