[论文解读] Architecture inside the mirage: evaluating generative image models on architectural style, elements, and typologies
研究评估五大 GenAI 图像平台在 30 条建筑提示上的表现,衡量生成图像对照历史学家标准的准确性,显示总体准确性有限,并提示需要标签化与出处溯源。
Generative artificial intelligence (GenAI) text-to-image systems are increasingly used to generate architectural imagery, yet their capacity to reproduce accurate images in a historically rule-bound field remains poorly characterized. We evaluated five widely used GenAI image platforms (Adobe Firefly, DALL-E 3, Google Imagen 3, Microsoft Image Generator, and Midjourney) using 30 architectural prompts spanning styles, typologies, and codified elements. Each prompt-generator pair produced four images (n = 600 images total). Two architectural historians independently scored each image for accuracy against predefined criteria, resolving disagreements by consensus. Set-level performance was summarized as zero to four accurate images per four-image set. Image output from Common prompts was 2.7-fold more accurate than from Rare prompts (p < 0.05). Across platforms, overall accuracy was limited (highest accuracy score 52 percent; lowest 32 percent; mean 42 percent). All-correct (4 out of 4) outcomes were similar across platforms. By contrast, all-incorrect (0 out of 4) outcomes varied substantially, with Imagen 3 exhibiting the fewest failures and Microsoft Image Generator exhibiting the highest number of failures. Qualitative review of the image dataset identified recurring patterns including over-embellishment, confusion between medieval styles and their later revivals, and misrepresentation of descriptive prompts (for example, egg-and-dart, banded column, pendentive). These findings support the need for visible labeling of GenAI synthetic content, provenance standards for future training datasets, and cautious educational use of GenAI architectural imagery.
研究动机与目标
- 评估五个广泛使用的 GenAI 图像平台在再现建筑风格、类型与文本提示要素方面的能力。
- 使用独立专家基于标准化标准对图像进行准确性评分并量化。
- 考察提示频率(常见 vs 罕见)对生成图像准确性的影响。
- 描述 GenAI 输出的定性模式以为标签化与出处标准提供信息。
提出的方法
- 使用五个平台:Adobe Firefly、DALL-E 3、Google Imagen 3、Microsoft Image Generator 与 Midjourney。
- 开发 30 条涵盖风格、类型与编码要素的建筑提示。
- 每个提示-平台组合生成四张图像(总计 n = 600 张图像)。
- 由两名建筑史学家独立对图像进行基于预定义标准的准确性评分;分歧通过共识解决。
- 按集合总结性能(每四张图像集合的4张中有多少张准确)。
- 对常见提示与罕见提示进行统计比较(p < 0.05)。
实验结果
研究问题
- RQ1五大 GenAI 平台在再现建筑风格、类型与要素方面的准确性水平如何?
- RQ2提示频率(常见 vs 罕见)如何影响输出准确性?
- RQ3在准确性和失误率方面是否存在平台特异性模式?
- RQ4在 GenAI 建筑图像中出现了哪些定性模式影响可靠性和可解释性?
主要发现
- 各平台的平均准确性为 42%(范围 32%–52%)。
- 常见提示的准确性比罕见提示高出 2.7 倍(p < 0.05)。
- 所观察到的最高准确性为 52%,最低为 32%;所有正确(4/4)的结果在各平台之间相似。
- 全错(0/4)结果因平台而异,Imagen 3 的失败最少,Microsoft Image Generator 的失败最多。
- 定性模式包括过度渲染、混淆中世纪风格与复兴风格,以及对描述性提示(如蛋-浮饰、带状柱、悬臂圆顶等)的错误表达。
- 研究结果支持对合成内容进行可见标签化与训练数据出处标准的要求;在教育场景中应谨慎使用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。