Skip to main content
QUICK REVIEW

[论文解读] Inferring Fine-grained Details on User Activities and Home Location from Social Media: Detecting Drinking-While-Tweeting Patterns in Communities

Nabil Hossain, Tianran Hu|arXiv (Cornell University)|Mar 10, 2016
Human Mobility and Location-Based Analysis参考文献 44被引用 20
一句话总结

该论文提出了一种机器学习框架,用于从推文数据中区分即时饮酒行为报告与过去/未来提及及一般性讨论,同时以街区级分辨率准确推断用户居住地。通过使用纽约市和门罗县的分层支持向量机分类器和地理标记推文,研究揭示了本地酒精销售点密度与社区层面饮酒自报数据之间存在显著正相关关系,且在郊区地区相关性更强。

ABSTRACT

Nearly all previous work on geo-locating latent states and activities from social media confounds general discussions about activities, self-reports of users participating in those activities at times in the past or future, and self-reports made at the immediate time and place the activity occurs. Activities, such as alcohol consumption, may occur at different places and types of places, and it is important not only to detect the local regions where these activities occur, but also to analyze the degree of participation in them by local residents. In this paper, we develop new machine learning based methods for fine-grained localization of activities and home locations from Twitter data. We apply these methods to discover and compare alcohol consumption patterns in a large urban area, New York City, and a more suburban and rural area, Monroe County. We find positive correlations between the rate of alcohol consumption reported among a community's Twitter users and the density of alcohol outlets, demonstrating that the degree of correlation varies significantly between urban and suburban areas. While our experiments are focused on alcohol use, our methods for locating homes and distinguishing temporally-specific self-reports are applicable to a broad range of behaviors and latent states.

研究动机与目标

  • 区分社交媒体中即时饮酒行为报告与过去/未来提及及一般性讨论。
  • 从稀疏、嘈杂的地理标记推文数据中,以街区级(100米分辨率)准确推断用户居住地。
  • 通过将即时报告与居住地及本地酒精销售点密度关联,分析社区层面的饮酒行为模式。
  • 比较城市(纽约市)与郊区(门罗县)社区的饮酒行为模式。
  • 利用社交媒体数据实现对酒精使用行为的细粒度公共卫生分析,克服传统调查的局限性。

提出的方法

  • 训练了三级支持向量机(SVM)分类器,将推文分类为:(1) 关于酒精的一般性讨论,(2) 过去/未来饮酒的自报,(3) 即时饮酒。
  • 使用人工标注的训练数据,基于时间具体性和自我指涉性对推文进行标注,使每个SVM分类器的F值均超过83%。
  • 开发了基于SVM的街区级居住地推断模型,仅需每位用户五条地理标记推文即可训练,实现100米网格内70%的准确率。
  • 将模型应用于纽约市和门罗县的地理标记推文数据,以比较城市与郊区的饮酒模式。
  • 将即时饮酒推文映射至用户居住地,并计算至饮酒地点的移动距离,以识别饮酒地点相对于居住地的位置。
  • 通过空间分析将社区层面的饮酒率与本地酒精销售点密度相关联,评估环境因素的影响。

实验结果

研究问题

  • RQ1如何在社交媒体中区分即时饮酒行为报告与过去/未来自报及一般性讨论?
  • RQ2在街区级分辨率下,多大程度上能从稀疏的地理标记推文数据中准确推断用户居住地?
  • RQ3社交媒体中反映的城市与郊区社区在饮酒行为模式上存在哪些差异?
  • RQ4酒精销售点密度与社区层面的Twitter自报饮酒行为之间存在何种相关性?
  • RQ5通过将居住地与即时饮酒推文关联,能获得关于移动模式和饮酒场所的哪些见解?

主要发现

  • 分层SVM分类器在区分即时饮酒报告与过去/未来提及及一般性讨论方面,F值均超过83%。
  • 居住地推断模型在100米网格内实现70%的准确率,覆盖了纽约市71%的活跃用户,且仅需每位用户五条地理标记推文。
  • 在社区层面,即时饮酒自报率与酒精销售点密度之间存在显著正相关关系,且在郊区地区相关性更强。
  • 在门罗县(一个以郊区为主的地区),酒精销售点密度与饮酒自报的相关性显著高于纽约市,表明在较不城市化的环境中环境影响更强。
  • 平均而言,报告即时饮酒的推文来自距离家超过1,000米的位置,表明饮酒通常不在家中进行。
  • 本研究证明,社交媒体数据可为社区层面的健康行为提供细粒度、实时的洞察,实现可扩展的公共卫生监测。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。