Skip to main content
QUICK REVIEW

[Paper Review] The emojification of sentiment on social media: Collection and analysis of a longitudinal Twitter sentiment dataset

Wenjie Yin, Rabab Alkhalifa|arXiv (Cornell University)|Aug 31, 2021
Sentiment Analysis and Opinion Mining14 references4 citations
TL;DR

This paper introduces TM-Senti, a large-scale, distantly supervised Twitter sentiment dataset of over 184 million tweets spanning seven years, labeled via emoticons and emojis. It reveals a longitudinal shift toward increased emoji usage in sentiment expression, with the dataset fully rehydratable using Internet Archive archives for reproducible research in sentiment analysis and text classification.

ABSTRACT

Social media, as a means for computer-mediated communication, has been extensively used to study the sentiment expressed by users around events or topics. There is however a gap in the longitudinal study of how sentiment evolved in social media over the years. To fill this gap, we develop TM-Senti, a new large-scale, distantly supervised Twitter sentiment dataset with over 184 million tweets and covering a time period of over seven years. We describe and assess our methodology to put together a large-scale, emoticon- and emoji-based labelled sentiment analysis dataset, along with an analysis of the resulting dataset. Our analysis highlights interesting temporal changes, among others in the increasing use of emojis over emoticons. We publicly release the dataset for further research in tasks including sentiment analysis and text classification of tweets. The dataset can be fully rehydrated including tweet metadata and without missing tweets thanks to the archive of tweets publicly available on the Internet Archive, which the dataset is based on.

Motivation & Objective

  • To address the lack of longitudinal sentiment datasets in social media research, particularly for tracking sentiment evolution over time.
  • To develop a large-scale, distantly supervised Twitter sentiment dataset using emoticons and emojis as proxy sentiment labels.
  • To analyze temporal trends in sentiment expression, especially the transition from emoticons to emojis in user-generated content.
  • To ensure dataset reproducibility and full rehydration by leveraging publicly archived tweets from the Internet Archive.
  • To support future research in sentiment analysis, text classification, and social media analytics through open data release.

Proposed method

  • The dataset was constructed using a distantly supervised approach, labeling tweets based on the presence of emoticons (e.g., `:)`, `:(`) and emojis (e.g., 😊, 😢) as sentiment proxies.
  • Over 184 million tweets were collected from the Internet Archive’s Twitter dataset, covering the period from 2013 to 2020.
  • Sentiment labels were assigned based on the sentiment polarity of the emoticons and emojis present in each tweet, with positive, negative, or neutral classification.
  • The dataset includes full tweet metadata (e.g., user ID, timestamp, retweet status) to support longitudinal and contextual analysis.
  • A validation process ensured label consistency and dataset quality, with corrections applied in later versions.
  • The dataset is fully rehydratable using the Internet Archive’s archived tweets, enabling full reconstruction and reproducibility.

Experimental results

Research questions

  • RQ1How has the usage of emojis and emoticons evolved in Twitter sentiment expression from 2013 to 2020?
  • RQ2To what extent do emojis serve as reliable sentiment indicators compared to traditional emoticons in social media texts?
  • RQ3What are the temporal trends in sentiment distribution across different topics or events in the dataset?
  • RQ4How does the shift from emoticons to emojis reflect broader changes in digital communication behavior?
  • RQ5Can a distantly supervised approach using emojis and emoticons produce a reliable and scalable sentiment dataset for longitudinal analysis?

Key findings

  • The dataset contains over 184 million tweets, with sentiment labels derived from emoticons and emojis, enabling large-scale sentiment analysis.
  • There is a significant and consistent increase in emoji usage relative to emoticons over the seven-year period, indicating a trend toward 'emojification' of sentiment.
  • The proportion of tweets containing emojis rose steadily from 2013 to 2020, while emoticon usage declined, reflecting a shift in digital expression.
  • The dataset is fully rehydratable using the Internet Archive’s tweet archive, ensuring reproducibility and access to original tweet metadata.
  • The distantly supervised labeling approach yielded a reliable and scalable dataset suitable for downstream tasks in sentiment analysis and text classification.
  • The dataset was corrected for a typo in the appendix across versions v1 to v3, with v3 released in March 2025.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.