[論文レビュー] Predicting and Understanding Turn-Taking Behavior in Open-Ended Group Activities in Virtual Reality
本論文は、動作、視線、性格特徴に基づく勾配ブースティングを用いて、VRのオープンエンドなグループ活動におけるターンテーキング挙動を予測し、0.71–0.78 AUCをwhat/who/whenタスク全体で達成し、顕著な特徴を特定する。
In networked virtual reality (VR), user behaviors, individual differences, and group dynamics can serve as important signals into future speech behaviors, such as who the next speaker will be and the timing of turn-taking behaviors. The ability to predict and understand these behaviors offers opportunities to provide adaptive and personalized assistance, for example helping users with varying sensory abilities navigate complex social scenes and instantiating virtual moderators with natural behaviors. In this work, we predict turn-taking behaviors using features extracted based on social dynamics literature. We discuss results from a large-scale VR classroom dataset consisting of 77 sessions and 1660 minutes of small-group social interactions collected over four weeks. In our evaluation, gradient boosting classifiers achieved the best performance, with accuracies of 0.71--0.78 AUC (area under the ROC curve) across three tasks concerning the "what", "who", and "when" of turn-taking behaviors. In interpreting these models, we found that group size, listener personality, speech-related behavior (e.g., time elapsed since the listener's last speech event), group gaze (e.g., how much the group looks at the speaker), as well as the listener's and previous speaker's head pitch, head y-axis position, and left hand y-axis position more saliently influenced predictions. Results suggested that these features remain reliable indicators in novel social VR settings, as prediction performance is robust over time and with groups and activities not used in the training dataset. We discuss theoretical and practical implications of the work.
研究の動機と目的
- 個人、グループ、動作/発話特徴から、VRのオープンエンドなグループ活動におけるターンテーキングを予測できるかを調査する。
- トレーニングで見られなかったグループ、活動、時間にわたるターンテーキング予測のロバスト性を評価する。
- ターンテーキング予測とモデル性能に最も影響を与える非言語的および人口統計的特徴を特定する。
提案手法
- 4週間にわたる77セッション、1660分のオープンエンドなグループディスカッションを含む大規模VR教室データセットを使用する。
- 30 Hzでの動作キャプチャと音声から、自己中心的な動作、二者間/グループ視線、対 interpersonal距離、頭部・手の姿勢を含む特徴を抽出する。
- 4つのターン遷移カテゴリー(クリーンターンタケ、オーバーラップ、バックチャネル、スピーチ継続)を定義し、IPUからターンをラベル付けする。
- 遷移前1秒の特徴ウィンドウを構築し、話者の直前10名の発話列と性格・グループ特徴をエンコードする。
- 次に話す人と時刻を予測する勾配ブースティング分類器を訓練し、AUCで評価する。
実験結果
リサーチクエスチョン
- RQ1RQ1: VRのオープンエンドなグループで、個人、グループ、発話、および動作特徴からターンテーキングを予測できるか。
- RQ2RQ2: トレーニングで見られなかったグループ、活動、時間へ予測性能はどれだけ転移するか。
- RQ3RQ3: ターンテーキング予測とモデル性能に最も関連する特徴は何か。
主な発見
- 勾配ブースティングは、次の話者を予測する際に0.75–0.78 AUC、ターン遷移の瞬間を予測する際に0.71–0.72 AUCという最良の精度を示した。
- 顕著な特徴には、リスナーの性格、グループサイズ、直前の発話列、リスナーの最後のターンからの経過時間、グループ視線、頭部ピッチ、頭部のy軸位置、左手のy軸位置が含まれる。
- トレーニング中に見られなかった時刻、活動、およびグループで評価した場合でも予測性能はロバストのままであった。
- 結果は、VRのソーシャル環境におけるターンテーキング予測に、特徴間の非線形相互作用が関与していることを示唆している。
- 知見は、没入型ソーシャル環境におけるリアルタイム介入と適応支援の理論的・実践的含意を提供する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。