[论文解读] Human Social Interaction Modeling Using Temporal Deep Networks
本文提出了一种新颖的混合深度学习模型——判别式条件受限玻尔兹曼机(DCRBM),以计算方式建模关键社会互动谓词(ESIPs),如共同注意和同步。通过联合检测ESIPs(准确率49–76%)并生成低层次行为数据(均方误差为0.01–0.1),该模型以数据驱动且可解释的方式揭示了社交亲密度的行为构成要素。
We present a novel approach to computational modeling of social interactions based on modeling of essential social interaction predicates (ESIPs) such as joint attention and entrainment. Based on sound social psychological theory and methodology, we collect a new "Tower Game" dataset consisting of audio-visual capture of dyadic interactions labeled with the ESIPs. We expect this dataset to provide a new avenue for research in computational social interaction modeling. We propose a novel joint Discriminative Conditional Restricted Boltzmann Machine (DCRBM) model that combines a discriminative component with the generative power of CRBMs. Such a combination enables us to uncover actionable constituents of the ESIPs in two steps. First, we train the DCRBM model on the labeled data and get accurate (76\%-49\% across various ESIPs) detection of the predicates. Second, we exploit the generative capability of DCRBMs to activate the trained model so as to generate the lower-level data corresponding to the specific ESIP that closely matches the actual training data (with mean square error 0.01-0.1 for generating 100 frames). We are thus able to decompose the ESIPs into their constituent actionable behaviors. Such a purely computational determination of how to establish an ESIP such as engagement is unprecedented.
研究动机与目标
- 开发一种计算框架,以识别和建模关键社会互动谓词(ESIPs),如共同注意和同步。
- 通过基于双人互动中经验性多模态数据的ESIP检测,弥合社会心理学与机器学习之间的鸿沟。
- 通过联合判别与生成建模,实现将ESIPs分解为可操作的行为构成要素。
- 创建一个新的基准数据集——积木游戏数据集,以支持计算社会互动建模的研究。
- 通过建模逼真的互动动态,支持跨文化培训、人机交互和虚拟现实等应用。
提出的方法
- 研究者使用佩戴在胸前的GoPro和Kinect传感器,采集在积木游戏中进行的双人互动的新型音视频数据集,并对ESIPs进行标注。
- 提出一种新颖的联合判别式条件受限玻尔兹曼机(DCRBM),结合判别分类与潜在表征的生成建模。
- DCRBM通过判别目标从多模态特征(骨骼、音频、视线)中检测ESIPs(如同步、模仿)。
- 检测完成后,模型利用其生成能力,基于ESIP类别标签合成低层次行为数据(如关节位置)。
- 通过计算生成序列与真实序列之间的归一化均方误差(NMSE),在不同长度(16–300帧)下评估生成性能。
- 在两种设置下评估模型:从部分输入生成缺失的玩家数据,以及仅从类别标签生成完整的可见层数据。
实验结果
研究问题
- RQ1如何从多模态传感器数据中计算检测关键社会互动谓词(ESIPs),如共同注意和同步?
- RQ2混合深度学习模型能否同时检测ESIPs并生成与训练数据匹配的逼真低层次行为数据?
- RQ3复杂社会谓词(如同步和模仿)背后的行为构成要素是什么?
- RQ4该模型在生成不同长度序列及不同强度ESIPs时的稳定性与准确性如何?
- RQ5DCRBM在多大程度上能够捕捉双人社会互动的语义与时间结构?
主要发现
- DCRBM在不同ESIPs上的检测准确率在49%至76%之间,证明了其对社会互动谓词的有效分类能力。
- 生成误差始终较低,大多数ESIPs的归一化均方误差低于0.1,表明生成逼真行为序列的高保真度。
- 模型在不同强度水平的ESIPs上均表现出稳定性能,表明对互动强度变化具有鲁棒性。
- 如预期,随着序列长度增加,生成误差上升,这是由于长时间段内真实数据方差增大所致。
- 模型仅从类别标签成功生成了完整的可见层数据(两名玩家),且误差较低,证明其在复杂社会动态生成方面的强大能力。
- 该框架实现了对ESIPs的纯计算分解为可操作的行为构成要素,这是计算社会科学研究中此前未见的能力。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。