[论文解读] Design Space and Implementation of RAG-Based Avatars for Virtual Archaeology
论文为虚拟现实中以文化遗产为主题的RAG驱动化身设计空间提供定义,并围绕麦森提乌斯陵墓构建VR原型,评估RAG配置与VR环境中的用户工作负载。
Immersive technologies, such as virtual and augmented reality, are transforming digital heritage by enabling users to explore and interact with culturally significant sites. It is now possible to view and augment digital twins, or digitally reconstructed versions of them, and to enable access to previously unreachable locations for a broader audience. Here, we investigate retrieval-augmented generation (RAG)-based avatars as an interface for accessing further information about digital cultural heritage objects while immersed in dedicated virtual environments. We present a requirement design space that spans the application realm, avatar personality, and I/O modalities. We instantiate it with a RAG system coupled to a conversational avatar in a virtual reality (VR) environment, using the Maxentius mausoleum from the 4th century AD as a case study, through which users gain access to curated on-demand information of the digitised heritage object. Our workflow utilises scholarly texts and enriches them with metadata. We evaluate various RAG configurations in terms of answer quality on a small expert-crafted question-answer set, as well as the perceived workload of users of a VR setup using such a RAG avatar. We demonstrate evidence that users perceive the overall workload for interacting with such an avatar as below average and that such avatars help to gain topical engagement. Overall, our work demonstrates how to utilise RAG-driven VR avatars for archaeological purposes and provides evidence that they can offer a pathway for immersive, AI-enhanced digital heritage applications.
研究动机与目标
- 为文化遗产应用的VR中AI驱动化身提出需求设计空间。
- 在VR环境中以麦森提乌斯陵墓为案例实例化基于RAG的化身原型。
- 评估不同RAG配置在VR交互中对答案质量与用户工作负载的影响。
- 展示沉浸式、AI增强数字遗产应用的可行性,并为从业者勾勒设计考虑因素。
提出的方法
- 开发一个三块式的概念设计(应用领域、化身个性、化身输入/输出)以指导系统需求。
- 设计一个包含本地数据存储、基于嵌入的检索和对话的TTS的逻辑VR–RAG架构。
- 实现一个基于Unity的物理VR原型,连接到本地FlowiseAI RAG堆栈,使用Qdrant向量存储和CIDOC-CRM知情元数据方法。
- curate a literature-based knowledge base focusing on the Maxentius mausoleum and enrich metadata with author, title, publication type, and relevance.
- 评估多种RAG配置(七种设置),使用专家生成的问答对和LLM作为评判标准,以及NASA-TLX工作量评估来评估VR用户。
实验结果
研究问题
- RQ1如何将AI驱动化身在VR文化遗产情境下的设计空间结构化,以覆盖应用、个性与输入/输出方面的考量?
- RQ2哪些RAG配置在本地VR部署的可计算性与答案质量之间提供最佳平衡?
- RQ3在VR中与基于RAG的化身互动如何影响用户工作负载和感知参与度?
- RQ4哪些实际工作流和元数据策略最能支持RAG-RV部署中的领域特定推理?
- RQ5一个RAG驱动的VR化身是否能够在保持受控创造性与准确性的前提下,有效提供关于特定考古对象(麦森提乌斯陵墓)的专家级信息?
主要发现
- RAG驱动的VR化身能够在VR文化遗产场景中提供按需的专家级信息。
- 不同的RAG配置会影响答案质量,具备领域感知的元数据与知识图谱能改进检索引导。
- 用户在与VR中的RAG化身互动时报告的工作负载偏低,表明在沉浸式遗产探索中具有良好的可用性。
- 纯本地端、低温度RAG设置结合 curated literature(精选文献)支持负责任、非创造性强度较低的回答,适用于学术用途。
- 麦森提乌斯陵墓案例展示了从OCR文本到CIDOC-CRM知情知识图谱驱动的GraphRAG系统的可行工作流。
- 研究提供证据表明AI驱动的化身可能促进沉浸式、AI增强的数字遗产应用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。