Skip to main content
QUICK REVIEW

[论文解读] Where Chicagoans tweet the most: Semantic analysis of preferential return locations of Twitter users

Aiman Soliman, Junjun Yin|arXiv (Cornell University)|Dec 21, 2015
Human Mobility and Location-Based Analysis参考文献 13被引用 4
一句话总结

本研究基于2014年来自164,000多名芝加哥居民的地理定位推文数据,识别出偏好返回位置及其语义用地类型。通过在语义增强的推文上应用DBSCAN聚类,发现尽管家和工作地点是常见的主要位置,但相当一部分用户偏离了这一模式,且发推行为与用地类型密切相关,强烈受时间限制,挑战了‘主要两个位置普遍为家和工作’的假设。

ABSTRACT

Recent studies on human mobility show that human movements are not random and tend to be clustered. In this connection, the movements of Twitter users captured by geo-located tweets were found to follow similar patterns, where a few geographic locations dominate the tweeting activity of individual users. However, little is known about the semantics (landuse types) and temporal tweeting behavior at those frequently-visited locations. Furthermore, it is generally assumed that the top two visited locations for most of the users are home and work locales (Hypothesis A) and people tend to tweet at their top locations during a particular time of the day (Hypothesis B). In this paper, we tested these two frequently cited hypotheses by examining the tweeting patterns of more than 164,000 unique Twitter users whom were residents of the city of Chicago during 2014. We extracted landuse attributes for each geo-located tweet from the detailed inventory of the Chicago Metropolitan Agency for Planning. Top-visited locations were identified by clustering semantic enriched tweets using a DBSCAN algorithm. Our results showed that although the top two locations are likely to be residential and occupational/educational, a portion of the users deviated from this case, suggesting that the first hypothesis oversimplify real-world situations. However, our observations indicated that people tweet at specific times and these temporal signatures are dependent on landuse types. We further discuss the implication of confounding variables, such as clustering algorithm parameters and relative accuracy of tweet coordinates, which are critical factors in any experimental design involving Twitter data.

研究动机与目标

  • 调查芝加哥推文用户的主要两个访问位置是否始终为家和工作(假设A)。
  • 检查高频访问位置的发推时间模式及其对用地类型依赖性的关系(假设B)。
  • 评估聚类算法参数和GPS精度等混杂变量对基于推文的移动性分析的影响。
  • 通过地理定位推文的语义增强,识别并分类偏好返回位置。
  • 通过社交媒体数据提供人类移动性语义与时间结构的实证证据。

提出的方法

  • 收集并处理了2014年芝加哥地区超过164,000名独立推文用户的地理定位推文数据。
  • 利用芝加哥大都会规划署的详细用地清单,为每条推文分配用地属性。
  • 应用DBSCAN聚类,基于空间邻近性和语义相似性对推文进行分组,以识别偏好返回位置。
  • 提取每个识别出位置的发推活动时间模式,以分析一天中不同时段的行为特征。
  • 通过评估DBSCAN参数和GPS坐标精度的敏感性,验证结果的稳健性。
  • 按用地类型(如居住、商业、教育等)对主要位置进行分类,以检验关于主导位置的假设。

实验结果

研究问题

  • RQ1大多数芝加哥推文用户的主要两个位置是否始终如一地为家和工作,如普遍假设的那样?
  • RQ2用户是否表现出与首选位置语义用地类型相关的显著时间性发推模式?
  • RQ3聚类算法参数和GPS精度在多大程度上影响基于推文数据识别偏好返回位置的结果?
  • RQ4用户返回非居住、非工作地点的频率如何?这些地点的语义特征是什么?
  • RQ5在居住区、教育区和商业区等不同用地类别之间,发推行为是否存在显著差异?

主要发现

  • 相当一部分推文用户并未将家和工作地点作为其前两位访问位置,挑战了‘家和工作普遍占主导地位’的常见假设。
  • 用户在发推行为上表现出强烈的时间特征,且与首选位置的语义用地类型密切相关。
  • 访问频率最高的位置主要为居住类和职业/教育类,但非传统地点(如休闲或商业区域)也显示出较高的访问频率。
  • 不同用地类型的发推时间模式存在显著差异:例如,工作地点在工作日白天达到发推高峰,而居住区则在晚间表现出更高活跃度。
  • 结果对聚类算法参数和GPS坐标精度敏感,凸显了在推文移动性研究中方法严谨性的重要性。
  • 对地理定位推文进行语义增强,可实现对超越原始空间聚类的更有意义的人类移动模式的精细化识别。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。