[Paper Review] Discovering Basic Emotion Sets via Semantic Clustering on a Twitter Corpus
This paper proposes a data-driven approach to discovering basic emotion sets using semantic clustering on a Twitter corpus, applying Latent Semantic Clustering (LSC) to evaluate semantic distinctiveness of emotion terms. It identifies a new, semantically more distinct emotion set—Accepting, Ashamed, Contempt, Interested, Joyful, Pleased, Sleepy, Stressed—showing a 6.1% improvement in distinctiveness over Ekman’s classic set.
A plethora of words are used to describe the spectrum of human emotions, but how many emotions are there really, and how do they interact? Over the past few decades, several theories of emotion have been proposed, each based around the existence of a set of 'basic emotions', and each supported by an extensive variety of research including studies in facial expression, ethology, neurology and physiology. Here we present research based on a theory that people transmit their understanding of emotions through the language they use surrounding emotion keywords. Using a labelled corpus of over 21,000 tweets, six of the basic emotion sets proposed in existing literature were analysed using Latent Semantic Clustering (LSC), evaluating the distinctiveness of the semantic meaning attached to the emotional label. We hypothesise that the more distinct the language is used to express a certain emotion, then the more distinct the perception (including proprioception) of that emotion is, and thus more 'basic'. This allows us to select the dimensions best representing the entire spectrum of emotion. We find that Ekman's set, arguably the most frequently used for classifying emotions, is in fact the most semantically distinct overall. Next, taking all analysed (that is, previously proposed) emotion terms into account, we determine the optimal semantically irreducible basic emotion set using an iterative LSC algorithm. Our newly-derived set (Accepting, Ashamed, Contempt, Interested, Joyful, Pleased, Sleepy, Stressed) generates a 6.1% increase in distinctiveness over Ekman's set (Angry, Disgusted, Joyful, Sad, Scared). We also demonstrate how using LSC data can help visualise emotions. We introduce the concept of an Emotion Profile and briefly analyse compound emotions both visually and mathematically.
Motivation & Objective
- To identify the most semantically distinct emotion sets by analyzing language use around emotion keywords on social media.
- To evaluate the psychological validity of existing emotion theories by measuring semantic distinctiveness through natural language data.
- To develop a method for discovering a semantically irreducible basic emotion set using iterative clustering on real-world linguistic data.
- To visualize emotional states using emotion profiles and multidimensional scaling, enabling analysis of compound emotions.
- To explore applications in clinical psychology, emotional engineering, and economic prediction through real-time emotion analysis.
Proposed method
- Constructed a labeled Twitter corpus of over 21,000 tweets using tracked emotion keywords and filtered phrases.
- Applied Latent Semantic Clustering (LSC) to analyze semantic similarity between emotion terms using cosine similarity of vector representations.
- Used Partial Singular Value Decomposition (SVD) to reduce dimensionality and extract latent semantic structures from the emotion word co-occurrence matrix.
- Implemented an iterative LSC algorithm to identify the optimal set of semantically irreducible emotions by maximizing distinctiveness.
- Developed an 'Emotion Profile' model using multidimensional scaling to visualize emotional states and compound emotions in a semantic space.
- Validated results using geographical relativity tests and compared LSC outputs against established emotion models like Ekman’s and Plutchik’s.
Experimental results
Research questions
- RQ1Which existing emotion set exhibits the highest semantic distinctiveness in natural language use on Twitter?
- RQ2Can a new, semantically more irreducible basic emotion set be discovered through data-driven clustering of real-world linguistic expressions?
- RQ3How can semantic clustering of emotion-related language be used to visualize and mathematically model compound emotional states?
- RQ4To what extent does the semantic distinctiveness of emotion terms correlate with their perceived psychological distinctiveness?
- RQ5Can LSC-based emotion analysis detect biases in public emotional expression, such as publicity bias, compared to private emotional language?
Key findings
- Ekman’s classic emotion set (Angry, Disgusted, Joyful, Sad, Scared) was found to be the most semantically distinct among the six tested emotion sets.
- The newly derived emotion set—Accepting, Ashamed, Contempt, Interested, Joyful, Pleased, Sleepy, Stressed—achieved a 6.1% increase in semantic distinctiveness over Ekman’s set.
- The LSC algorithm successfully identified a semantically irreducible set of eight emotions that better represent the full spectrum of human emotional experience.
- Emotion Profiles generated via multidimensional scaling effectively visualized both individual and compound emotional states, such as 'depressed' or 'guilty'.
- The analysis revealed that combinations of primary emotions (e.g., Joyful + Scared) showed measurable similarity to compound states like 'depressed', supporting mathematical modeling of emotional blends.
- The study demonstrated that semantic clustering of social media language can detect patterns of emotional expression with potential applications in clinical assessment and economic prediction.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.