Skip to main content
QUICK REVIEW

[Paper Review] Inferring Fine-grained Details on User Activities and Home Location from Social Media: Detecting Drinking-While-Tweeting Patterns in Communities

Nabil Hossain, Tianran Hu|arXiv (Cornell University)|Mar 10, 2016
Human Mobility and Location-Based Analysis44 references20 citations
TL;DR

This paper proposes a machine learning framework to distinguish in-the-moment alcohol consumption reports from past/future mentions and general discussions in Twitter data, while accurately inferring user home locations at block-level resolution. Using hierarchical SVM classifiers and geo-tagged tweets from New York City and Monroe County, it reveals a significant positive correlation between local alcohol outlet density and community-level drinking self-reports, with stronger correlations in suburban areas.

ABSTRACT

Nearly all previous work on geo-locating latent states and activities from social media confounds general discussions about activities, self-reports of users participating in those activities at times in the past or future, and self-reports made at the immediate time and place the activity occurs. Activities, such as alcohol consumption, may occur at different places and types of places, and it is important not only to detect the local regions where these activities occur, but also to analyze the degree of participation in them by local residents. In this paper, we develop new machine learning based methods for fine-grained localization of activities and home locations from Twitter data. We apply these methods to discover and compare alcohol consumption patterns in a large urban area, New York City, and a more suburban and rural area, Monroe County. We find positive correlations between the rate of alcohol consumption reported among a community's Twitter users and the density of alcohol outlets, demonstrating that the degree of correlation varies significantly between urban and suburban areas. While our experiments are focused on alcohol use, our methods for locating homes and distinguishing temporally-specific self-reports are applicable to a broad range of behaviors and latent states.

Motivation & Objective

  • To distinguish in-the-moment alcohol consumption reports from past/future mentions and general discussions in social media.
  • To accurately infer user home locations from sparse, noisy geo-tagged Twitter data at block-level (100m resolution).
  • To analyze community-level alcohol consumption patterns by linking in-the-moment reports to home locations and local alcohol outlet density.
  • To compare drinking behavior patterns between urban (New York City) and suburban (Monroe County) communities.
  • To enable fine-grained public health analysis of alcohol use using social media data, overcoming limitations of traditional surveys.

Proposed method

  • Trained a hierarchy of three support vector machines (SVMs) to classify tweets into: (1) general discussion of alcohol, (2) self-reports of past/future drinking, and (3) in-the-moment drinking.
  • Used human-annotated training data to label tweets based on temporal specificity and self-reference, achieving F-scores above 83% for each SVM classifier.
  • Developed a block-level home location inference model using SVMs trained on as few as five geo-tagged tweets per user, achieving 70% accuracy within 100m grids.
  • Applied the models to geo-tagged Twitter data from New York City and Monroe County to compare urban and suburban drinking patterns.
  • Mapped in-the-moment drinking tweets to home locations and calculated travel distances to identify drinking locations relative to home.
  • Correlated community-level drinking rates with local alcohol outlet density using spatial analysis to assess environmental influences.

Experimental results

Research questions

  • RQ1How can we distinguish in-the-moment alcohol consumption reports from past/future self-reports and general discussions in social media?
  • RQ2To what extent can user home locations be accurately inferred from sparse geo-tagged Twitter data at block-level resolution?
  • RQ3What are the differences in alcohol consumption patterns between urban and suburban communities as reflected in social media?
  • RQ4How does the density of alcohol outlets correlate with community-level self-reported drinking on Twitter?
  • RQ5What insights can be gained about mobility patterns and drinking settings by linking home locations to in-the-moment drinking tweets?

Key findings

  • The hierarchical SVM classifiers achieved F-scores above 83% in distinguishing in-the-moment drinking reports from past/future mentions and general discussions.
  • The home location inference model achieved 70% accuracy in predicting user homes within 100m grids, covering 71% of active users in New York City with as few as five geo-tagged tweets.
  • A significant positive correlation was found between the rate of in-the-moment drinking self-reports and the density of alcohol outlets in the community, with stronger correlations observed in suburban areas.
  • In Monroe County, a predominantly suburban area, the correlation between outlet density and drinking self-reports was significantly higher than in New York City, suggesting stronger environmental influence in less urban settings.
  • On average, tweets reporting in-the-moment drinking were sent from locations more than 1,000 meters from home, indicating that drinking is often not done at home.
  • The study demonstrates that social media data can provide fine-grained, real-time insights into community-level health behaviors, enabling scalable public health monitoring.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.