Skip to main content
QUICK REVIEW

[Paper Review] Word Embedding for Social Sciences: An Interdisciplinary Survey

Akira Matsui, Emilio Ferrara|arXiv (Cornell University)|Jul 7, 2022
Computational and Text Analysis Methods4 citations
TL;DR

This survey provides a comprehensive taxonomy of word embedding applications in social sciences, focusing on word2vec and related models to extract semantic representations from text and non-text data. It highlights methodological trends, warns about inconsistencies in similarity measurements, and demonstrates how word embeddings enable variable construction across qualitative and quantitative research.

ABSTRACT

To extract essential information from complex data, computer scientists have been developing machine learning models that learn low-dimensional representation mode. From such advances in machine learning research, not only computer scientists but also social scientists have benefited and advanced their research because human behavior or social phenomena lies in complex data. However, this emerging trend is not well documented because different social science fields rarely cover each other's work, resulting in fragmented knowledge in the literature. To document this emerging trend, we survey recent studies that apply word embedding techniques to human behavior mining. We built a taxonomy to illustrate the methods and procedures used in the surveyed papers, aiding social science researchers in contextualizing their research within the literature on word embedding applications. This survey also conducts a simple experiment to warn that common similarity measurements used in the literature could yield different results even if they return consistent results at an aggregate level.

Motivation & Objective

  • To document the growing trend of applying word embedding techniques to social science research, particularly in analyzing unstructured or unconventional data.
  • To address the lack of methodological clarity in social science applications by building a standardized taxonomy of word embedding usage.
  • To identify and discuss emerging applications beyond text, including non-textual human behavior data such as online community interactions.
  • To caution researchers about potential inconsistencies in similarity measurements used in word embedding analysis, even when aggregate results appear consistent.
  • To bridge the gap between computer science and social science by clarifying how word2vec and related models can be effectively and rigorously applied in interdisciplinary research.

Proposed method

  • Developed a nine-label taxonomy to categorize how social scientists apply word embedding models, focusing on variable definition, reference words, and theoretical grounding.
  • Reviewed 27 journal publications and working papers from diverse social science fields, classifying them according to the taxonomy based on methodological choices.
  • Utilized pre-trained word2vec models (e.g., CBOW, skip-gram with negative sampling) as the primary tool, emphasizing their use in low-dimensional semantic representation of words.
  • Conducted a simple experiment to compare common similarity metrics (e.g., cosine similarity, Euclidean distance) and demonstrate their divergent results at the individual level despite consistent aggregate outcomes.
  • Applied word embedding techniques to infer latent variables such as partisanship, sentiment, and demographic characteristics from online communities (e.g., Reddit subreddits) without relying on human-annotated labels.
  • Emphasized the use of reference words (e.g., 'democrats' and 'conservatives') to define working variables in non-textual contexts, enabling data-driven, linguistically independent variable construction.

Experimental results

Research questions

  • RQ1How are word embedding models, particularly word2vec, being applied across diverse social science disciplines beyond traditional text analysis?
  • RQ2What methodological patterns emerge in the use of word embeddings for constructing research variables in social science studies?
  • RQ3To what extent do common similarity metrics used in word embedding analysis yield consistent results at the individual level, despite consistent aggregate performance?
  • RQ4How can word embedding models be adapted to infer social characteristics (e.g., partisanship, gender, age) from non-textual or indirect behavioral data in online communities?
  • RQ5What are the key challenges and limitations in applying word embeddings to qualitative and interdisciplinary social science research?

Key findings

  • Word embedding models are increasingly used across social sciences—not only in quantitative fields like finance and management science but also in qualitative domains such as history, cultural studies, and political science.
  • A significant trend is the use of word embeddings to define research variables indirectly through reference words (e.g., 'democrats' and 'conservatives' on Reddit), enabling data-driven, linguistically independent variable construction.
  • The study reveals that similarity metrics such as cosine and Euclidean distance can produce divergent results at the individual level, even when aggregate-level results appear consistent, highlighting a critical methodological risk.
  • Applications extend beyond text to infer social characteristics from non-textual behavioral data, such as user activity patterns in online communities, demonstrating the versatility of word embeddings in capturing social phenomena.
  • The taxonomy developed in this survey identifies recurring methodological patterns, including the use of pre-trained models, reference-based variable definition, and clustering or prediction tasks, which can guide future research design.
  • Despite widespread use, many social science papers do not specify whether they use CBOW or skip-gram models, indicating a need for greater methodological transparency in reporting word embedding applications.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.