[论文解读] A Domain Based Approach to Social Relation Recognition
本文提出一种基于领域的社会关系识别方法,利用社会心理学中的五个领域(依恋、互惠、求偶、等级权力和结盟群体)构建分层标签空间。研究引入了一个包含26,915对人物标注的新数据集,并表明基于语义属性的模型优于其他方法,在性能和可解释性方面均表现出色,且与心理学理论保持一致。
Social relations are the foundation of human daily life. Developing techniques to analyze such relations from visual data bears great potential to build machines that better understand us and are capable of interacting with us at a social level. Previous investigations have remained partial due to the overwhelming diversity and complexity of the topic and consequently have only focused on a handful of social relations. In this paper, we argue that the domain-based theory from social psychology is a great starting point to systematically approach this problem. The theory provides coverage of all aspects of social relations and equally is concrete and predictive about the visual attributes and behaviors defining the relations included in each domain. We provide the first dataset built on this holistic conceptualization of social life that is composed of a hierarchical label space of social domains and social relations. We also contribute the first models to recognize such domains and relations and find superior performance for attribute based features. Beyond the encouraging performance of the attribute based approach, we also find interpretable features that are in accordance with the predictions from social psychology literature. Beyond our findings, we believe that our contributions more tightly interleave visual recognition and social psychology theory that has the potential to complement the theoretical work in the area with empirical and data-driven models of social life.
研究动机与目标
- 开发一个全面的计算框架,用于识别视觉数据中的社会关系,涵盖人类社会生活的各个方面。
- 通过基于整体心理学理论的方法,解决以往研究仅关注少数临时社会关系类别的局限性。
- 构建一个大规模、分层的数据集,包含社会领域和关系级别的标注,以提升泛化能力和评估效果。
- 评估从社会心理学中提取的语义属性在从图像中识别社会领域和关系方面的有效性。
- 通过构建可解释、理论对齐的模型,弥合理论社会心理学与数据驱动视觉识别之间的鸿沟。
提出的方法
- 将社会心理学中的布吉塔尔领域理论(Bugental’s domain-based theory)改编为概念框架,将社会关系划分为五个领域:依恋、互惠、求偶、等级权力和结盟群体。
- 通过为图像中的人物对分配领域级和关系级标签,构建分层标签空间,基于扩展后的PIPA数据集(新增26,915个标注)。
- 设计并训练基于语义属性的模型,使用年龄、性别、服装、活动以及身体/头部外观等属性,这些属性基于心理学预测进行设计。
- 训练并比较端到端的全数据驱动模型与基于属性的模型,以评估性能和可解释性。
- 从外部数据集迁移预训练的属性模型(如服装和活动)以提升特征学习和泛化能力。
- 进行消融研究和归一化贡献分析,以评估各类属性在分类领域和关系中的相对重要性。
实验结果
研究问题
- RQ1社会心理学中的基于领域的理论能否作为建模视觉数据中社会关系的全面且具有预测力的框架?
- RQ2基于属性的模型在识别社会领域和关系方面,与端到端深度学习模型相比表现如何?
- RQ3所学习的视觉属性在多大程度上与社会心理学文献中的预测一致?
- RQ4哪些视觉属性对分类社会领域和关系最具预测力,且其贡献在不同任务中如何变化?
- RQ5通过学习领域级表征,模型能否在不同社会关系之间实现泛化?
主要发现
- 基于语义属性的模型在识别社会领域和关系方面均优于全数据驱动模型,表明理论启发的特征更具有效性。
- 对领域和关系识别贡献最高的属性是活动和服装,这与基于领域的理论强调社会群体中共享行为和外貌的观点一致。
- 年龄和性别特征(尤其是头部年龄和身体性别)是重要预测因子,尤其在求偶和依恋领域中表现显著,证实了社会心理学中的预测。
- 即使在图像质量较低或背景遮挡的困难情况下,模型性能依然稳健,例如在雾霾或部分可见的人物对中仍能做出正确预测。
- 负样本通常涉及异常行为或模糊线索(例如,祖母爬行以抱住婴儿),尽管个体属性信号强烈,仍会使模型产生混淆。
- 当视觉线索不清晰时(如新闻发布会的低分辨率图像),模型表现下降,凸显了在受限视觉条件下改进外观建模的必要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。