Skip to main content
QUICK REVIEW

[Paper Review] Sharing emotions at scale: The Vent dataset

Nikolaos Lykousas, Constantinos Patsakis|arXiv (Cornell University)|Jan 15, 2019
Complex Network Analysis Techniques4 citations
TL;DR

This paper introduces the Vent dataset, the largest publicly available collection of user-generated text with fine-grained emotion annotations, comprising 33 million posts from nearly 1 million users across 705 distinct emotions organized into 63 categories. The dataset enables large-scale research in affective computing by providing self-reported emotion labels, revealing patterns in emotional expression, temporal dynamics, and social network structures, with strong alignment to existing emotion lexicons like EmoLex.

ABSTRACT

The continuous and increasing use of social media has enabled the expression of human thoughts, opinions, and everyday actions publicly at an unprecedented scale. We present the Vent dataset, the largest annotated dataset of text, emotions, and social connections to date. It comprises more than 33 millions of posts by nearly a million of users together with their social connections. Each post has an associated emotion. There are 705 different emotions, organized in 63 "emotion categories", forming a two-level taxonomy of affects. Our initial statistical analysis describes the global patterns of activity in the Vent platform, revealing large heterogenities and certain remarkable regularities regarding the use of the different emotions. We focus on the aggregated use of emotions, the temporal activity, and the social network of users, and outline possible methods to infer emotion networks based on the user activity. We also analyze the text and describe the affective landscape of Vent, finding agreements with existing (small scale) annotated corpus in terms of emotion categories and positive/negative valences. Finally, we discuss possible research questions that can be addressed from this unique dataset.

Motivation & Objective

  • To address the scarcity of large-scale, high-quality emotion-annotated datasets in affective computing.
  • To provide a comprehensive resource for studying emotional expression, temporal patterns, and social network dynamics in user-generated content.
  • To enable the development and evaluation of advanced emotion recognition models, particularly deep learning approaches.
  • To explore emotional homophily and contagion in online social networks using real-world user data.
  • To establish a benchmark for emotion analysis beyond binary sentiment, capturing the full spectrum of human affect.

Proposed method

  • The Vent dataset was collected from the Vent social networking app, where users voluntarily share emotional posts and self-assign emotions from a predefined set of 705 emotion labels.
  • Emotions are organized into a two-level taxonomy with 63 emotion categories, enabling hierarchical analysis of affective expression.
  • Textual content was processed and analyzed for valence using the EmoLex lexicon, allowing comparison with established affective resources.
  • Statistical and network analysis were applied to study temporal activity patterns, emotional distribution, and social connectivity among users.
  • The dataset includes social graph data, including user follow relationships and interaction patterns (e.g., reactions like 'hug', 'same', 'h4u').
  • Neural network-based models are proposed as a methodological pathway to leverage the dataset’s scale and diversity for emotion recognition.

Experimental results

Research questions

  • RQ1How do emotional expressions vary across different emotion categories in large-scale social media content?
  • RQ2To what extent do users exhibit emotional homophily—forming connections with others expressing similar emotions?
  • RQ3Can emotional contagion be detected and modeled in online social networks using real user data?
  • RQ4How do the valence scores of Vent emotion categories compare to established affective lexicons like EmoLex?
  • RQ5What are the temporal patterns and regularities in emotional expression across different user groups and emotion types?

Key findings

  • The Vent dataset contains 33 million posts from nearly 1 million users, with 705 distinct emotion labels organized into 63 emotion categories, forming a comprehensive two-level taxonomy of affect.
  • The distribution of valence scores for 'Happiness' and 'Affection' categories shows a clear positive skew, aligning with their positive emotional valence, while other categories are predominantly negative or neutral.
  • There is strong agreement between the Vent emotion categories and the EmoLex annotated corpus, particularly in the alignment of positive and negative valence groupings.
  • Temporal analysis reveals significant heterogeneity in user activity, with notable regularities in the use of specific emotions across different times of day and days of the week.
  • The social network structure shows evidence of emotional homophily, with users tending to follow and interact with others expressing similar emotions.
  • The dataset enables the study of emotional contagion, with potential to model how emotions spread through social links, supported by the availability of both user activity and network structure data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.