[논문 리뷰] Architecture inside the mirage: evaluating generative image models on architectural style, elements, and typologies
이 연구는 30개 건축 프롬프트에 대해 다섯 개의 GenAI 이미지 플랫폼을 평가하고, 생성 이미지의 정확도를 역사학자 기준에 따라 측정했으며, 전반적인 정확도가 제한적임을 밝히고 표식 및 출처 표기에 대한 필요성을 제기한다.
Generative artificial intelligence (GenAI) text-to-image systems are increasingly used to generate architectural imagery, yet their capacity to reproduce accurate images in a historically rule-bound field remains poorly characterized. We evaluated five widely used GenAI image platforms (Adobe Firefly, DALL-E 3, Google Imagen 3, Microsoft Image Generator, and Midjourney) using 30 architectural prompts spanning styles, typologies, and codified elements. Each prompt-generator pair produced four images (n = 600 images total). Two architectural historians independently scored each image for accuracy against predefined criteria, resolving disagreements by consensus. Set-level performance was summarized as zero to four accurate images per four-image set. Image output from Common prompts was 2.7-fold more accurate than from Rare prompts (p < 0.05). Across platforms, overall accuracy was limited (highest accuracy score 52 percent; lowest 32 percent; mean 42 percent). All-correct (4 out of 4) outcomes were similar across platforms. By contrast, all-incorrect (0 out of 4) outcomes varied substantially, with Imagen 3 exhibiting the fewest failures and Microsoft Image Generator exhibiting the highest number of failures. Qualitative review of the image dataset identified recurring patterns including over-embellishment, confusion between medieval styles and their later revivals, and misrepresentation of descriptive prompts (for example, egg-and-dart, banded column, pendentive). These findings support the need for visible labeling of GenAI synthetic content, provenance standards for future training datasets, and cautious educational use of GenAI architectural imagery.
연구 동기 및 목표
- 다섯 가지 널리 사용되는 GenAI 이미지 플랫폼이 텍스트 프롬프트로부터 건축 양식, 유형, 요소를 얼마나 잘 재현하는지 평가한다.
- 표준화된 기준에 따른 독립 전문가 점수를 사용하여 이미지 정확도를 정량화한다.
- 프롬프트 빈도(일반 vs 희귀)가 생성 이미지의 정확도에 미치는 영향을 검토한다.
- 라벨링 및 출처 표준화를 위한 GenAI 산출물의 질적 패턴을 특성화한다.
제안 방법
- 다섯 가지 GenAI 플랫폼을 사용한다: Adobe Firefly, DALL-E 3, Google Imagen 3, Microsoft Image Generator, 그리고 Midjourney.
- 스타일, 유형, 그리고 체계화된 요소를 포괄하는 30개의 건축 프롬프트를 개발한다.
- 프롬프트-플랫폼 쌍당 네 장의 이미지를 생성한다(총 600장).
- 두 명의 건축사 역사가 독립적으로 사전에 정해진 기준에 따라 정확도를 점수화하도록 한다; 이견은 합의로 해결한다.
- 세트당 성과를 요약한다(네 이미지 세트당 0~4개 정확한 이미지).
- 일반 프롬프트와 희귀 프롬프트의 통계적 비교(p < 0.05).
실험 결과
연구 질문
- RQ1건축 양식, 유형, 요소를 재현하는 데 있어 GenAI 플랫폼 전반의 정확도 수준은 어느 정도인가?
- RQ2프롬프트 빈도(일반 vs 희귀)가 출력 정확도에 어떤 영향을 미치는가?
- RQ3정확도와 실패율에 플랫폼별 패턴이 있는가?
- RQ4신뢰성과 해석 가능성에 영향을 미치는 GenAI 건축 이미지의 질적 패턴은 무엇인가?
주요 결과
- 플랫폼 간 평균 정확도는 42%(범위 32%–52%).
- 일반 프롬프트가 희귀 프롬프트보다 정확도가 2.7배 높다(p < 0.05).
- 관찰된 최고 정확도는 52%, 최저는 32%; 전부 정확(4/4) 결과는 플랫폼 간 비슷하다.
- 전부 잘못된(0/4) 결과는 플랫폼에 따라 다르며, Imagen 3이 실패가 가장 적고 Microsoft Image Generator가 가장 많다.
- 질적 패턴으로는 과도한 장식, 중세 양식과 부흥 양식의 혼동, 설명 프롬프트의 오해(예: 에그-앤-다트, 띠 모양 기둥, 팬던티브) 등이 포함된다.
- 발견은 합성 콘텐츠의 가시적 표식 및 교육 데이터의 출처 표준화에 대한 필요성을 지지하며, 교육 용도에 신중한 사용을 권고한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.