Skip to main content
QUICK REVIEW

[Paper Review] Architecture inside the mirage: evaluating generative image models on architectural style, elements, and typologies

Jamie Magrill, Leah Gornstein|arXiv (Cornell University)|Jan 14, 2026
Aesthetic Perception and Analysis0 citations
TL;DR

The study evaluates five GenAI image platforms on 30 architectural prompts, measuring accuracy of generated images against historian-criteria, revealing limited overall accuracy and prompting labeling and provenance needs.

ABSTRACT

Generative artificial intelligence (GenAI) text-to-image systems are increasingly used to generate architectural imagery, yet their capacity to reproduce accurate images in a historically rule-bound field remains poorly characterized. We evaluated five widely used GenAI image platforms (Adobe Firefly, DALL-E 3, Google Imagen 3, Microsoft Image Generator, and Midjourney) using 30 architectural prompts spanning styles, typologies, and codified elements. Each prompt-generator pair produced four images (n = 600 images total). Two architectural historians independently scored each image for accuracy against predefined criteria, resolving disagreements by consensus. Set-level performance was summarized as zero to four accurate images per four-image set. Image output from Common prompts was 2.7-fold more accurate than from Rare prompts (p < 0.05). Across platforms, overall accuracy was limited (highest accuracy score 52 percent; lowest 32 percent; mean 42 percent). All-correct (4 out of 4) outcomes were similar across platforms. By contrast, all-incorrect (0 out of 4) outcomes varied substantially, with Imagen 3 exhibiting the fewest failures and Microsoft Image Generator exhibiting the highest number of failures. Qualitative review of the image dataset identified recurring patterns including over-embellishment, confusion between medieval styles and their later revivals, and misrepresentation of descriptive prompts (for example, egg-and-dart, banded column, pendentive). These findings support the need for visible labeling of GenAI synthetic content, provenance standards for future training datasets, and cautious educational use of GenAI architectural imagery.

Motivation & Objective

  • Assess how well five widely used GenAI image platforms reproduce architectural styles, typologies, and elements from textual prompts.
  • Quantify image accuracy using independent expert scoring across standardized criteria.
  • Examine the impact of prompt frequency (Common vs Rare) on generated image accuracy.
  • Characterize qualitative patterns in GenAI outputs to inform labeling and provenance standards.

Proposed method

  • Use five GenAI platforms: Adobe Firefly, DALL-E 3, Google Imagen 3, Microsoft Image Generator, and Midjourney.
  • Develop 30 architectural prompts spanning styles, typologies, and codified elements.
  • Generate four images per prompt-platform pair (n = 600 images).
  • Have two architectural historians independently score images for accuracy against predefined criteria; resolve disagreements by consensus.
  • Summarize performance per set (0–4 accurate images per four-image set).
  • Statistical comparison of Common vs Rare prompts (p < 0.05).

Experimental results

Research questions

  • RQ1What is the level of accuracy across GenAI platforms in reproducing architectural styles, typologies, and elements?
  • RQ2How does prompt frequency (Common vs Rare) influence output accuracy?
  • RQ3Are there platform-specific patterns in accuracy and failure rates?
  • RQ4What qualitative patterns emerge in GenAI architectural imagery that affect reliability and interpretability?

Key findings

  • Mean accuracy across platforms is 42% (range 32%–52%).
  • Common prompts yield 2.7x higher accuracy than Rare prompts (p < 0.05).
  • Highest accuracy observed is 52%, lowest 32%; all-correct (4/4) outcomes are similar across platforms.
  • All-incorrect (0/4) outcomes vary by platform, with Imagen 3 showing the fewest failures and Microsoft Image Generator the most.
  • Qualitative patterns include over-embellishment, confusion between medieval styles and revivals, and misrepresentation of descriptive prompts (e.g., egg-and-dart, banded column, pendentive).
  • Findings support visible labeling of synthetic content and provenance standards for training data; advise cautious use in education.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.