Skip to main content
QUICK REVIEW

[논문 리뷰] Predicting and Understanding Turn-Taking Behavior in Open-Ended Group Activities in Virtual Reality

Portia Wang, Eugy Han|arXiv (Cornell University)|2024. 07. 03.
Team Dynamics and Performance인용 수 6
한 줄 요약

이 논문은 모션, 시선, 성격 특성에 대한 그래디언트 부스팅을 이용해 열린 그룹 활동에서 VR의 차례를 예측하고, 무엇/누구/언제 과제에서 0.71–0.78 AUC를 달성하며 중요한 특징들을 식별한다.

ABSTRACT

In networked virtual reality (VR), user behaviors, individual differences, and group dynamics can serve as important signals into future speech behaviors, such as who the next speaker will be and the timing of turn-taking behaviors. The ability to predict and understand these behaviors offers opportunities to provide adaptive and personalized assistance, for example helping users with varying sensory abilities navigate complex social scenes and instantiating virtual moderators with natural behaviors. In this work, we predict turn-taking behaviors using features extracted based on social dynamics literature. We discuss results from a large-scale VR classroom dataset consisting of 77 sessions and 1660 minutes of small-group social interactions collected over four weeks. In our evaluation, gradient boosting classifiers achieved the best performance, with accuracies of 0.71--0.78 AUC (area under the ROC curve) across three tasks concerning the "what", "who", and "when" of turn-taking behaviors. In interpreting these models, we found that group size, listener personality, speech-related behavior (e.g., time elapsed since the listener's last speech event), group gaze (e.g., how much the group looks at the speaker), as well as the listener's and previous speaker's head pitch, head y-axis position, and left hand y-axis position more saliently influenced predictions. Results suggested that these features remain reliable indicators in novel social VR settings, as prediction performance is robust over time and with groups and activities not used in the training dataset. We discuss theoretical and practical implications of the work.

연구 동기 및 목표

  • 개방형 VR 그룹 활동에서 차례가 지켜지는지 여부를 개인, 그룹, 모션/발화 특징으로 예측할 수 있는지 조사한다.
  • 학습 중에 보지 못한 그룹, 활동, 시간에서도 차례 예측의 강건성을 평가한다.
  • 차례 예측 및 모델 성능에 가장 영향을 미치는 비언어적 비상구? 비언어적 및 인구통계적 특징을 식별한다.

제안 방법

  • 4주에 걸친 77개 세션, 1660분의 개방형 그룹 토론을 포함하는 대규모 VR 교실 데이터세트를 사용한다.
  • 모션 캡처와 오디오에서 초당 30프레임으로 에고센트릭 모션, 이인 및 그룹 시선, 대인 간 거리, 머리/손 자세를 포함한 특징을 추출한다.
  • 네 가지 차례 전환 카테고리(정돈된 차례, 겹침, 백채널링, 발화 지속)를 정의하고 IPU에서 차례를 라벨링한다.
  • 전환 1초 전의 특징 윈도우를 구성하고 발화 시퀀스(이전 10명의 화자)와 성격 및 그룹 특징을 인코딩한다.
  • 다음에 누가 발화할지와 언제 발화할지 예측하도록 그래디언트 부스팅 분류기를 학습하고 AUC로 평가한다.

실험 결과

연구 질문

  • RQ1RQ1: VR 개방형 그룹에서 차례가 개인, 그룹, 발화 및 모션 특징으로 예측될 수 있는가?
  • RQ2RQ2: 학습에 보이지 않은 그룹, 활동, 시간에 예측 성능이 어떻게 전달되는가?
  • RQ3RQ3: 차례 예측 및 모델 성능과 가장 관련된 특징은 무엇인가?

주요 결과

  • 그래디언트 부스팅은 다음 화자를 예측하는 데 0.75–0.78 AUC, 차례 전환 순간을 예측하는 데 0.71–0.72 AUC의 최고 정확도를 달성했다.
  • 주요 특징으로는 청취자 성격, 그룹 크기, 선행 발화 시퀀스, 청취자의 마지막 차례 이후 경과 시간, 그룹 시선, 머리 피치, 머리 축 y 위치, 왼손 축 y 위치가 포함된다.
  • 학습 중 보지 못한 시간, 활동, 그룹에서 평가했을 때 예측 성능이 강건하게 유지되었다.
  • 특징 간의 비선형 상호작용이 VR 사회적 환경에서 차례 예측의 기저를 이룬다는 것을 시사한다.
  • 실시간 개입 및 몰입형 사회 환경에서의 적응적 지원에 대한 이론적·실용적 시사점을 제공한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.