Skip to main content
QUICK REVIEW

[논문 리뷰] Investigating the Design Considerations for Integrating Text-to-Image Generative AI within Augmented Reality Environments

Yongquan Hu, Zhang, Dawen|arXiv (Cornell University)|2023. 03. 29.
Augmented Reality Applications인용 수 4
한 줄 요약

이 논문은 음성 입력을 통해 AI가 생성한 이미지와 텍스트를 세 가지 AR 모odal리티—공간형, 헤드마운트형, 휴대형—에 투영하는 프로토타입인 GenerativeAIR를 구현한, 텍스트에서 이미지로의 생성형 AI를 증강현실(AR) 디스플레이와 통합하기 위한 설계 프레임워크를 제안한다. 주요 기여는 포커스 그룹 분석을 통해 유저, 기능, 환경을 통합한 설계 사고 모델을 도출한 것으로, AIGC+AR 시스템에서 개인정보 보호, 상호작용, 다중 사용자 협업과 같은 핵심 설계 고려사항을 규명하였다.

ABSTRACT

Generative Artificial Intelligence (GenAI) has emerged as a fundamental component of intelligent interactive systems, enabling the automatic generation of multimodal media content. The continuous enhancement in the quality of Artificial Intelligence-Generated Content (AIGC), including but not limited to images and text, is forging new paradigms for its application, particularly within the domain of Augmented Reality (AR). Nevertheless, the application of GenAI within the AR design process remains opaque. This paper aims to articulate a design space encapsulating a series of criteria and a prototypical process to aid practitioners in assessing the aptness of adopting pertinent technologies. The proposed model has been formulated based on a synthesis of design insights garnered from ten experts, obtained through focus group interviews. Leveraging these initial insights, we delineate potential applications of GenAI in AR.

연구 동기 및 목표

  • 텍스트에서 이미지로의 생성형 AI를 AR 디스플레이와 통합하기 위한 설계 공간을 탐색하기 위해.
  • 실세계 응용에서 AI 생성 콘텐츠(AIGC)와 AR를 통합할 때의 통합 설계 가이던스 부족 문제를 해결하기 위해.
  • AIGC+AR 시스템 설계에 영향을 미치는 사용자 수요, 기능적 요구사항, 환경적 요인를 규명하기 위해.
  • 세 가지 AR 디스플레이 유형에서 음성 기반의 다중 모odal AIGC 출력을 가능하게 하는 프로토타입을 개발하기 위해.
  • 실증적 포커스 그룹 연구를 통해 '사용자-기능-환경' 설계 사고 모델을 도출하기 위해.

제안 방법

  • Google의 음성-텍스트 변환 API를 스테이블 디퓨전 및 텍스트-음성 변환 모델과 통합하여 음성 기반 콘텐츠 생성을 구현한 프로토타입인 GenerativeAIR를 개발하였다.
  • 삼성 프리스타일 프로젝터(공간형 AR), 헬로우스 2(HMD), 옵티머스 10 프로(휴대형 디스플레이) 등 세 가지 AR 디스플레이 유형에 시스템을 구축하였다.
  • 다양한 시나리오에서 프로토타입을 사용한 참가자들이 참여한 포커스 그룹을 통해 사용자 피드백을 수집하였다.
  • 사용자 수요, 특히 개인정보 보호, 상호작용, 다중 사용자 조율과 관련된 반복적인 주제를 규명하기 위해 정성적 분석을 적용하였다.
  • AIGC+AR 시스템 설계 고려사항을 체계화하기 위해 '사용자-기능-환경' 설계 사고 모델을 개발하였다.
  • 실내/실외 환경과 사용자 역할(프레젠테이션 vs. 관찰자)에 따라 시스템 성능과 사용자 경험을 평가하였다.
Figure 1. The workflow and display effect of our GenerativeAIR prototype: (a) The user’s speech into the microphone is converted into text, which is then fed into an AI model for generating artistic images and more text; (b) Generated content in Spatial Augmented Reality (SAR): an example of Samsung
Figure 1. The workflow and display effect of our GenerativeAIR prototype: (a) The user’s speech into the microphone is converted into text, which is then fed into an AI model for generating artistic images and more text; (b) Generated content in Spatial Augmented Reality (SAR): an example of Samsung

실험 결과

연구 질문

  • RQ1텍스트에서 이미지로의 생성형 AI는 공간형, HMD, 휴대형 등 다양한 AR 디스플레이 모달리티에 어떻게 효과적으로 통합되어 실시간으로 상호작용 가능한 콘텐츠 생성을 지원할 수 있는가?
  • RQ2AI 생성 콘텐츠를 AR에 구현할 때 고려해야 할 설계 요소—특히 사용자 역할, 개인정보 보호, 환경적 맥락과 관련된 요소—는 무엇인가?
  • RQ3사용자들은 AI 생성 콘텐츠와 AR 디스플레이 간의 상호작용을 몰입도, 사용성, 사회적 상호작용 측면에서 어떻게 인지하는가?
  • RQ4생성형 AI를 사용할 때 단일 사용자와 다중 사용자 AR 경험 간의 기능적 및 경험적 차이는 무엇인가?
  • RQ5AIGC+AR 시스템에서 계층적 액세스와 콘텐츠 공유를 어떻게 구현할 수 있을까? 이는 개인정보 보호와 협업의 균형을 유지하기 위해 필수적이다.

주요 결과

  • 사용자들은 개인 정보(예: 사진, 일상 일기 등)를 포함한 AI 생성 콘텐츠의 경우, 개인정보 보호와 맥락 인식을 우선시한다.
  • 휴대형 디스플레이(예: 스마트폰)는 이동성과 접근성 덕분에 실시간 창작 미디어 생성에 가장 넓은 잠재력을 지닌다.
  • 공간형 AR 디스플레이는 공개적이고 공동의 경험을 강화하지만, 콘텐츠가 공개적으로 노출됨으로써 개인정보 보호 문제가 발생할 수 있다.
  • 프레젠테이션 사용자(예: HMD 사용자)는 제어력과 개인정보 보호를 중시하는 반면, 관찰자 사용자는 몰입도와 낮은 인지 부하를 우선시한다.
  • '사용자-기능-환경' 설계 모델은 역할 기반 액세스와 환경 적응과 같은 AIGC+AR 시스템 설계의 핵심 트레이드오프를 효과적으로 반영한다.
  • 현재 프로토타입는 2D 이미지 생성과 오프라인 운영에 국한되어 있어, 향후 연구에서는 실시간 3D 콘텐츠 생성과 온라인 통합이 필요하다.
Figure 2. The comparisons of AR display+generative AI and their related techniques: (a) display performance comparison of AR, VR and normal monitor; (b) content-generation performance comparison of generative AI (machine), AI assist (machine+human) and human.
Figure 2. The comparisons of AR display+generative AI and their related techniques: (a) display performance comparison of AR, VR and normal monitor; (b) content-generation performance comparison of generative AI (machine), AI assist (machine+human) and human.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.