[论文解读] Community Question Answering Platforms vs. Twitter for Predicting Characteristics of Urban Neighbourhoods
本文研究了来自 Yahoo! Answers(一个社区问答(QA)平台)的文本是否能够预测伦敦社区的真实世界人口统计特征,并与 Twitter 微博内容进行对比。通过使用自然语言处理(NLP)和回归模型分析用户生成的文本,研究发现 QA 文本的预测相关性平均为 ρ = 0.54,略高于 Twitter 的 ρ = 0.53;同时揭示了两种平台之间存在显著不同的语义模式:QA 文本体现百科全书式知识,而 Twitter 文本则反映当前的社会文化趋势。
In this paper, we investigate whether text from a Community Question Answering (QA) platform can be used to predict and describe real-world attributes. We experiment with predicting a wide range of 62 demographic attributes for neighbourhoods of London. We use the text from QA platform of Yahoo! Answers and compare our results to the ones obtained from Twitter microblogs. Outcomes show that the correlation between the predicted demographic attributes using text from Yahoo! Answers discussions and the observed demographic attributes can reach an average Pearson correlation coefficient of {ho} = 0.54, slightly higher than the predictions obtained using Twitter data. Our qualitative analysis indicates that there is semantic relatedness between the highest correlated terms extracted from both datasets and their relative demographic attributes. Furthermore, the correlations highlight the different natures of the information contained in Yahoo! Answers and Twitter. While the former seems to offer a more encyclopedic content, the latter provides information related to the current sociocultural aspects or phenomena.
研究动机与目标
- 评估非地理标记的、来自类似 Yahoo! Answers 的社区问答平台的文本是否能够预测城市社区的真实世界人口统计特征。
- 对比 Yahoo! Answers 文本与 Twitter 微博在预测伦敦社区 62 个人口统计特征方面的预测性能。
- 分析预测模型中高系数词语与对应人口统计特征之间的语义关系。
- 对比问答平台(百科全书式)与社交媒体(当前社会文化趋势)在城市属性建模中的信息本质差异。
- 证明尽管缺乏地理位置信息,QA 文本仍能为城市规划和社会学分析提供强有力的预测信号。
提出的方法
- 使用标准自然语言处理技术(包括分词、停用词去除和词形还原)对来自 Yahoo! Answers 中关于伦敦社区的讨论文本进行收集与预处理。
- 训练逻辑回归模型,利用从 QA 和 Twitter 文本中提取的 TF-IDF 加权特征,预测 62 个人口统计特征。
- 通过 10 折交叉验证评估模型性能,以皮尔逊相关系数(ρ)为主要评价指标。
- 提取预测模型中绝对系数最高的词语,以分析其与人口统计特征的语义相关性。
- 通过定性分析对比两个平台中最具相关性的词语及其对应的人口统计特征,突出内容类型的差异。
- 使用 Wilcoxon 符号秩检验评估 Yahoo! Answers 与 Twitter 在性能差异上的统计显著性。
实验结果
研究问题
- RQ1来自类似 Yahoo! Answers 的社区问答平台的文本是否能够以与 Twitter 相当或更优的性能,预测城市社区的真实世界人口统计特征?
- RQ2Yahoo! Answers 和 Twitter 文本中高影响力词语的语义模式如何与它们所预测的人口统计特征相关联?
- RQ3问答平台与微博平台在描述城市社区时所捕捉的信息本质有何不同?
- RQ4QA 数据缺乏地理位置信息是否会削弱其对社区层面属性的预测能力?
- RQ5某些人口统计特征(如种族、就业情况、年龄组)是否在某一平台上比另一平台更容易被预测?
主要发现
- 使用 Yahoo! Answers 文本预测 62 个人口统计特征的平均皮尔逊相关系数(ρ)为 0.54,略高于 Twitter 的 0.53。
- Yahoo! Answers 在预测与种族和就业相关的属性方面优于 Twitter,而 Twitter 在预测年龄组和汽车拥有率方面表现更佳。
- Wilcoxon 符号秩检验确认了两个平台之间的性能差异具有统计显著性(p < 0.01)。
- 定性分析显示,两个平台中最具预测力的词语与其对应的人口统计特征之间表现出强烈的语义一致性,例如在宗教人口统计讨论中,“Jewish”(犹太人)和“Jewish communities”(犹太社区)等词语。
- Yahoo! Answers 的文本特征更偏向百科全书式和事实性内容,而 Twitter 文本则反映实时的社会文化情绪和当前事件。
- 尽管 Yahoo! Answers 缺乏地理位置数据,其文本仍能为社区层面的人口统计特征提供稳健的预测信号。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。