Skip to main content
QUICK REVIEW

[论文解读] Sharing emotions at scale: The Vent dataset

Nikolaos Lykousas, Constantinos Patsakis|arXiv (Cornell University)|Jan 15, 2019
Complex Network Analysis Techniques被引用 4
一句话总结

本论文介绍了Vent数据集,这是目前公开可用的最大规模用户生成文本数据集,具备细粒度情感标注,包含近一百万用户发布的3300万条帖子,涵盖705种不同情绪,组织为63个类别。该数据集通过提供用户自报的情感标签,揭示了情感表达模式、时间动态及社交网络结构,为大规模情感计算研究提供了支持,且与现有情感词典(如EmoLex)具有高度一致性。

ABSTRACT

The continuous and increasing use of social media has enabled the expression of human thoughts, opinions, and everyday actions publicly at an unprecedented scale. We present the Vent dataset, the largest annotated dataset of text, emotions, and social connections to date. It comprises more than 33 millions of posts by nearly a million of users together with their social connections. Each post has an associated emotion. There are 705 different emotions, organized in 63 "emotion categories", forming a two-level taxonomy of affects. Our initial statistical analysis describes the global patterns of activity in the Vent platform, revealing large heterogenities and certain remarkable regularities regarding the use of the different emotions. We focus on the aggregated use of emotions, the temporal activity, and the social network of users, and outline possible methods to infer emotion networks based on the user activity. We also analyze the text and describe the affective landscape of Vent, finding agreements with existing (small scale) annotated corpus in terms of emotion categories and positive/negative valences. Finally, we discuss possible research questions that can be addressed from this unique dataset.

研究动机与目标

  • 解决情感计算领域中大规模、高质量情感标注数据集稀缺的问题。
  • 提供一个全面的资源,用于研究用户生成内容中的情感表达、时间模式及社交网络动态。
  • 推动先进情感识别模型(尤其是深度学习方法)的开发与评估。
  • 利用真实用户数据,探索在线社交网络中的情感同质性与情感传染现象。
  • 建立超越二元情感分析的基准,捕捉人类情感的完整光谱。

提出的方法

  • Vent数据集来源于Vent社交应用,用户自愿分享情感类帖子,并从预设的705种情绪标签中自行标注情绪。
  • 情绪被组织为两级分类体系,包含63个情绪类别,支持情感表达的分层分析。
  • 对文本内容使用EmoLex词典进行效价分析,使结果可与现有情感资源进行对比。
  • 应用统计分析与网络分析方法,研究用户活动的时间模式、情感分布及用户间社交连接性。
  • 数据集包含社交图谱数据,包括用户关注关系及互动模式(如‘拥抱’、‘同感’、‘h4u’等反应)。
  • 提出基于神经网络的模型作为方法路径,以利用该数据集的规模与多样性,实现情感识别。

实验结果

研究问题

  • RQ1在大规模社交媒体内容中,不同情绪类别的情感表达有何差异?
  • RQ2用户在多大程度上表现出情感同质性——即是否倾向于与表达相似情绪的人建立联系?
  • RQ3能否利用真实用户数据,在线社交网络中检测并建模情感传染现象?
  • RQ4Vent情绪类别的效价评分与EmoLex等既定情感词典相比有何异同?
  • RQ5不同用户群体与情绪类型之间,情感表达的时间模式与规律性如何?

主要发现

  • Vent数据集包含近一百万用户发布的3300万条帖子,涵盖705种不同情绪标签,组织为63个情绪类别,形成全面的两级情感分类体系。
  • ‘Happiness’(快乐)与‘Affection’(喜爱)类别的效价评分分布呈现明显的正偏态,与它们的正向情感效价一致;而其他类别的效价则主要为负向或中性。
  • Vent情绪类别与EmoLex标注语料库之间存在高度一致性,尤其在正负效价分组的对齐方面表现突出。
  • 时间分析揭示了用户活动的显著异质性,且在一天中不同时间段及一周内不同日期,特定情绪的使用存在明显规律。
  • 社交网络结构显示出情感同质性的证据,用户倾向于关注并互动于表达相似情绪的其他用户。
  • 该数据集支持情感传染的研究,具备建模情绪通过社交关系传播的潜力,得益于用户活动数据与网络结构数据的双重可及性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。