Skip to main content
QUICK REVIEW

[论文解读] A Large-Scale Study of Personal Identifiability of Virtual Reality Motion Over Time

Mark Roman Miller, Eugy Han|arXiv (Cornell University)|Mar 2, 2023
Virtual Reality Applications and Impacts被引用 11
一句话总结

本研究分析在随时间的虚拟现实(VR)运动数据可识别性,涵盖232名参与者、八周的多次会话,结果显示更多会话和每次会话的更长时长会提高可识别性,而更长的训练-测试延迟会降低可识别性;此外,本文引入了体空间坐标并评估多分类AUC。

ABSTRACT

In recent years, social virtual reality (VR), sometimes described as the "metaverse," has become widely available. With its potential comes risks, including risks to privacy. To understand these risks, we study the identifiability of participants' motion in VR in a dataset of 232 VR users with eight weekly sessions of about thirty minutes each, totaling 764 hours of social interaction. The sample is unique as we are able to study the effect of user, session, and time independently. We find that the number of sessions recorded greatly increases identifiability, and duration per session increases identifiability as well, but to a lesser degree. We also find that greater delay between training and testing sessions reduces identifiability. Ultimately, understanding the identifiability of VR activities will help designers, security professionals, and consumer advocates make VR safer.

研究动机与目标

  • 评估VR运动数据的可识别性如何随时间和不同数据规模演变。
  • 量化会话数量、单次会话时长以及训练与测试之间的时间延迟如何影响可识别性。
  • 引入体空间坐标以改善跨用户的特征对齐。
  • 展示可从运动数据推断人口属性(如性别和种族)的能力。
  • 建立评估指标(多分类AUC),以便在不同数据集和类别规模之间进行比较。

提出的方法

  • 使用斯坦福纵向VR课堂数据集,包含232名参与者,分两次采集、八次每周约30分钟的会话。
  • 在使用ENGAGE平台进行社交VR讨论时记录头戴设备与手持控制器的位置与旋转。
  • 从42个流(位置、旋转及派生的相对运动)中构造840个特征,包括体空间坐标。
  • 将体空间坐标系统定义为相对于每个参与者的前进方向,以提高对水平方向旋转的不变性。
  • 训练一个随机森林分类器(R/ranger),由600棵树组成的集成,并对每个会话的预测进行聚合以实现多类识别。
  • 以多类AUC作为主要评估指标,辅以在固定N分类测试集上的准确率。
Figure 1: Parallel coordinates plot of classification size, span of time in which data was collected, and total duration of data collected per participant. The current work is the largest or the second-largest on all dimensions. Note all dimensions are log-scaled in order to better scale the variati
Figure 1: Parallel coordinates plot of classification size, span of time in which data was collected, and total duration of data collected per participant. The current work is the largest or the second-largest on all dimensions. Note all dimensions are log-scaled in order to better scale the variati

实验结果

研究问题

  • RQ1随着训练数据增加(更多会话)与单次会话时间变长,可识别性如何变化?
  • RQ2训练与测试数据之间的时间延迟如何在跨周的情境中影响可识别性?
  • RQ3将数据转换为体空间坐标是否优于全局坐标在可识别性方面?
  • RQ4是否能够从VR运动数据中推断性别、种族等人口属性,准确率如何?

主要发现

  • 可识别性随着记录的会话增多显著提高,且在一定程度上随着会话时长增加而提升。
  • 训练与测试数据之间的延迟在所研究的时间尺度上降低可识别性。
  • 在同一会话内的可识别性高于跨不同会话之间,与先前工作一致。
  • 本研究推荐多类AUC作为在不同类别规模数据集之间比较可识别性的稳健度量。
  • 体空间坐标有助于将运动特征对齐到用户的前进方向,从而改进识别用的特征集合。
  • 人口属性(性别和种族)可以从运动数据中推断,相较基线模型有小到中等的提升。
Figure 2: Participants performing discussion activities in the VR environment. In the top left panel, participants illustrate the environmental impacts of an oil spill with a duck model covered in black smudges representing oil. In the top right panel, several students discuss the experience of the
Figure 2: Participants performing discussion activities in the VR environment. In the top left panel, participants illustrate the environmental impacts of an oil spill with a duck model covered in black smudges representing oil. In the top right panel, several students discuss the experience of the

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。