[论文解读] Predicting and Understanding Turn-Taking Behavior in Open-Ended Group Activities in Virtual Reality
该论文在开放式群体活动中,使用基于运动、凝视和性格特征的梯度提升来预测虚拟现实中的轮流发言行为,在 what/who/when 任务中实现0.71–0.78 AUC,并识别出显著特征。
In networked virtual reality (VR), user behaviors, individual differences, and group dynamics can serve as important signals into future speech behaviors, such as who the next speaker will be and the timing of turn-taking behaviors. The ability to predict and understand these behaviors offers opportunities to provide adaptive and personalized assistance, for example helping users with varying sensory abilities navigate complex social scenes and instantiating virtual moderators with natural behaviors. In this work, we predict turn-taking behaviors using features extracted based on social dynamics literature. We discuss results from a large-scale VR classroom dataset consisting of 77 sessions and 1660 minutes of small-group social interactions collected over four weeks. In our evaluation, gradient boosting classifiers achieved the best performance, with accuracies of 0.71--0.78 AUC (area under the ROC curve) across three tasks concerning the "what", "who", and "when" of turn-taking behaviors. In interpreting these models, we found that group size, listener personality, speech-related behavior (e.g., time elapsed since the listener's last speech event), group gaze (e.g., how much the group looks at the speaker), as well as the listener's and previous speaker's head pitch, head y-axis position, and left hand y-axis position more saliently influenced predictions. Results suggested that these features remain reliable indicators in novel social VR settings, as prediction performance is robust over time and with groups and activities not used in the training dataset. We discuss theoretical and practical implications of the work.
研究动机与目标
- Investigate whether turn-taking in open-ended VR group activities can be predicted from individual, group, and motion/verbal features.
- Assess robustness of turn-taking predictions across groups, activities, and time not seen during training.
- Identify which nonverbal and demographic features most influence turn-taking predictions and model performance.
提出的方法
- Use a large-scale VR classroom dataset with 77 sessions, 1660 minutes of open-ended group discussions, across four weeks.
- Extract features from motion capture and audio at 30 Hz, including egocentric motion, dyadic/group gaze, interpersonal distances, and head/hand poses.
- Define four turn-transition categories (clean turn taking, overlap, backchanneling, continuing speech) and label turns from IPUs.
- Construct 1-second pre-transition feature windows and encode speech sequences (10 preceding speakers) plus personality and group features.
- Train gradient boosting classifiers to predict who will speak next and when, evaluating using AUC.
实验结果
研究问题
- RQ1RQ1: Can turn-taking in VR open-ended groups be predicted from individual, group, speech, and motion features?
- RQ2RQ2: How does prediction performance transfer to groups, activities, and times not seen in training?
- RQ3RQ3: Which features are most related to turn-taking predictions and model performance?
主要发现
- Gradient boosting yielded best accuracies with 0.75–0.78 AUC for predicting the next speaker and 0.71–0.72 AUC for predicting the turn-transition moment.
- Salient features include listener personality, group size, preceding speech sequences, time since listener’s last turn, group gaze, head pitch, head y-axis position, and left-hand y-axis position.
- Predictive performance remained robust when evaluated on time, activities, and groups not seen during training.
- Results suggest non-linear interactions among features underlie turn-taking predictions in VR social settings.
- Findings offer theoretical and practical implications for real-time interventions and adaptive assistance in immersive social environments.
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。