Skip to main content
QUICK REVIEW

[论文解读] Twitter Activity Timeline as a Signature of Urban Neighborhood

Philipp Kats, Cheng Qian|arXiv (Cornell University)|Jul 19, 2017
Human Mobility and Location-Based Analysis参考文献 32被引用 3
一句话总结

本文提出将推特活动时间序列(TWS)用作动态、数据驱动的签名,以识别纽约市的城市社区功能与社会经济特征。通过分析按邮政编码聚合的地理标记推文,作者构建了每日活动模式的时间序列档案,用于检测事件、基于功能相似性对社区进行聚类,并利用机器学习模型预测社会经济指标,R²最高达0.65。

ABSTRACT

Modern cities are complex systems, evolving at a fast pace. Thus, many urban planning, political, and economic decisions require a deep and up-to-date understanding of the local context of urban neighborhoods. This study shows that the structure of openly available social media records, such as Twitter, offers a possibility for building a unique dynamic signature of urban neighborhood function, and, therefore, might be used as an efficient and simple decision support tool. Considering New York City as an example, we investigate how Twitter data can be used to decompose the urban landscape into self-defining zones, aligned with the functional properties of individual neighborhoods and their social and economic characteristics. We further explore the potential of these data for detecting events and evaluating their impact over time and space. This approach paves a way to a methodology for immediate quantification of the impact of urban development programs and the estimation of socioeconomic statistics at a finer spatial-temporal scale, thus allowing urban policy-makers to track neighborhood transformations and foresee undesirable changes in order to take early action before official statistics would be available.

研究动机与目标

  • 开发一种利用实时、公开的推特数据量化城市社区功能的方法。
  • 基于社交媒体活动的时间模式,识别并表征具有不同功能特征的城市区域。
  • 利用推特时间序列检测异常事件(例如暴风雪),并量化其时空影响。
  • 利用TWS特征在邮政编码层面建模社会经济指标(例如通勤时间、教育水平、房价)。
  • 评估在整合特定应用来源的推文时,数据粒度与模型性能之间的权衡。

提出的方法

  • 从2014年1月至2016年6月,在纽约市四个行政区(排除史泰登岛)收集地理标记推文数据,排除自动化账户。
  • 将推文聚合至262个纽约市邮政编码,为每个区域创建每日活动模式的时间序列表示(TWS)。
  • 将TWS定义为168维向量(24小时×7天),总结每日每小时的平均推文数量。
  • 通过分离主要应用(原生推特、Instagram、Foursquare)提升TWS的应用粒度,将特征空间扩展至2,688维。
  • 对TWS应用聚类(k-means)以基于活动节律识别功能相似的社区。
  • 在TWS特征上训练监督机器学习模型(随机森林、额外树),以预测社会经济变量。

实验结果

研究问题

  • RQ1推特活动时间序列(TWS)能否作为城市社区功能的可靠且动态的表征?
  • RQ2在具有不同社会经济与土地利用特征的社区中,TWS模式有何差异?
  • RQ3在多大程度上可利用TWS检测并量化极端天气等城市事件的时空影响?
  • RQ4TWS特征能否以足够高的准确度预测社会经济指标(如教育水平、房价),以支持城市政策决策?
  • RQ5在使用细粒度应用级推文数据时,模型性能与数据需求之间的权衡如何?

主要发现

  • TWS模式成功识别出七个对应于不同功能城市区域的邮政编码聚类,包括商业核心、住宅区和混合用途区。
  • 使用应用特定TWS(2,688个特征)的模型在预测商业用地比例方面达到R²为0.62,优于基线TWS模型。
  • 额外树回归模型在预测高价住房单位比例方面达到R²为0.65,表明对房地产特征具有强大的预测能力。
  • 随机森林模型在预测研究生学位普及率方面达到R²为0.55,在预测大学文凭比例方面达到R²为0.54,表明对教育水平的估计具有可靠性。
  • 事件检测成功识别出一场重大暴风雪,并估算出各社区的暴露程度,显示出在应急响应和影响评估中的实用性。
  • 包含应用级特征的增强模型需要更多数据,且因稀疏性限制在96个邮政编码内,凸显了分辨率与数据可得性之间的权衡。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。