Skip to main content
QUICK REVIEW

[Paper Review] From Prompts to Worlds: How Users Iterate, Explore, and Make Sense of AI-Generated 3D Environments

Aung Pyae|arXiv (Cornell University)|Jan 24, 2026
Social Robot Interaction and HRI0 citations
TL;DR

This study empirically investigates how users interact with a commercial text-to-3D platform, revealing a language-to-space gap, episodic presence, and iteration barriers that shape sensemaking in AI-created 3D environments.

ABSTRACT

Text-to-3D generative AI systems create navigable environments from natural language prompts, but unlike text-to-image generation, evaluation requires embodied exploration of spatial coherence, scale, and navigability. We present the first empirical study of a commercial text-to-3D platform, combining think-aloud protocols, behavioral observation, and validated measures of usability, presence, and engagement. We report three findings. First, asymmetric expressibility: users readily convey semantic intent (themes, atmosphere) but struggle to specify spatial structure (layout, scale), reflecting a language-to-space limitation rather than a skill deficit. Second, episodic presence: immersion arises when expectations align with outputs but does not accumulate into sustained place illusion. Third, structural iteration breakdowns: refinement fails due to interaction barriers - poor discoverability, opaque feedback, and high temporal costs - rather than user limitations. Together, these dynamics form a reinforcing cycle in which spatial mismatches persist, producing episodic presence and ongoing sensemaking. We reframe text-to-3D interaction as negotiated meaning-making rather than linear prompting, and argue that effective systems require hybrid input modalities, transparent feedback, and low-cost iteration.

Motivation & Objective

  • Motivate understanding of how users translate natural language prompts into navigable 3D spaces.
  • Examine how users iterate, explore, and make sense of AI-generated 3D environments through embodied tasks.
  • Identify cognitive and interaction barriers that affect usability, presence, and engagement in text-to-3D systems.
  • Propose design implications for hybrid input modalities, transparent feedback, and low-cost iteration to improve user experience.

Proposed method

  • Combine think-aloud protocols with behavioral observation during interaction with a commercial text-to-3D platform.
  • Use validated measures of usability, presence, and engagement to assess user experience.
  • Analyze how semantic intent is expressed versus spatial structure specification to identify language-to-space limitations.
  • Characterize episodes of presence and their relation to expectation alignment with outputs.
  • Identify points of refinement breakdown and their underlying causes, such as discoverability and feedback opacity.

Experimental results

Research questions

  • RQ1How do users express semantic intent and spatial structure when using text-to-3D prompts?
  • RQ2What are the patterns of presence and immersion in AI-generated 3D environments, and how do they relate to alignment between expectations and outputs?
  • RQ3What interaction barriers impede systematic refinement and iteration in text-to-3D tools?
  • RQ4What design changes could mitigate language-to-space gaps and support lower-cost, more transparent iteration?
  • RQ5How should text-to-3D systems be framed in terms of meaning-making rather than linear prompting?

Key findings

  • Users readily convey semantic themes and atmosphere but struggle with specifying spatial layout and scale.
  • Immersion (episodic presence) arises when outputs align with expectations but does not accumulate into sustained place illusion.
  • Refinement breakdowns stem from interaction barriers such as poor discoverability, opaque feedback, and high temporal costs.
  • A reinforcing cycle emerges where spatial mismatches persist, leading to episodic presence and ongoing sensemaking.
  • The study argues for viewing text-to-3D interaction as negotiated meaning-making and suggests hybrid inputs, transparent feedback, and low-cost iteration.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.