[论文解读] What Twitter Data Tell Us about the Future?
本研究利用LDA与BERTopic对来自Twitter未来学家的逾100万条推文进行分析,以建模预期中的未来,LDA识别出15个主题,BERTopic则识别出100个不同的主题。研究揭示了‘未来现在’型未来——即动态、演化的潜在未来概念化形式——影响社交媒体用户主动预见并采取行动。结果表明,有影响力人物的语言线索可通过可扩展的NLP管道与开源数据及代码,塑造集体的前瞻性行为。
Anticipation is a fundamental human cognitive ability that involves thinking about and living towards the future. While language markers reflect anticipatory thinking, research on anticipation from the perspective of natural language processing is limited. This study aims to investigate the futures projected by futurists on Twitter and explore the impact of language cues on anticipatory thinking among social media users. We address the research questions of what futures Twitter's futurists anticipate and share, and how these anticipated futures can be modeled from social data. To investigate this, we review related works on anticipation, discuss the influence of language markers and prestigious individuals on anticipatory thinking, and present a taxonomy system categorizing futures into "present futures" and "future present". This research presents a compiled dataset of over 1 million publicly shared tweets by future influencers and develops a scalable NLP pipeline using SOTA models. The study identifies 15 topics from the LDA approach and 100 distinct topics from the BERTopic approach within the futurists' tweets. These findings contribute to the research on topic modelling and provide insights into the futures anticipated by Twitter's futurists. The research demonstrates the futurists' language cues signals futures-in-the-making that enhance social media users to anticipate their own scenarios and respond to them in present. The fully open-sourced dataset, interactive analysis, and reproducible source code are available for further exploration.
研究动机与目标
- 探究基于Twitter的未来学家所预测的未来类型及其对公众前瞻性思维的影响。
- 考察语言标记与知名人物在塑造社交媒体用户对未来情景预期方面的作用。
- 开发一种可扩展的NLP管道,用于从社交媒体数据中建模预期中的未来。
- 为未来研究提供公开可获取、可复现的数据集与分析框架,以支持前瞻性修辞与社会预测研究。
提出的方法
- 通过网络抓取,在Twitter学术API的限制与数据可得性考量下,收集了超过100万条来自认证Twitter未来学家的公开推文。
- 采用标准NLP技术对文本数据进行预处理,包括清洗、分词与噪声去除,为主题建模做准备。
- 应用潜在狄利克雷分布(LDA)进行主题建模,以识别未来学家推文中15个广泛的主题聚类。
- 采用BERTopic——一种基于Transformer的主题建模方法——从同一数据集中提取出100个更细致且语义连贯的主题。
- 开发了一种可扩展的、符合SOTA标准的NLP管道,专为短文本社交媒体内容设计,利用上下文嵌入与聚类技术。
- 通过开放科学框架(Open Science Framework)发布完整数据集、交互式可视化与源代码,以确保可复现性与社区再利用。

实验结果
研究问题
- RQ1RQ.1:Twitter上的未来学家预见并分享了哪些未来?
- RQ2RQ.2:如何利用NLP技术从社交媒体数据中建模预期中的未来?
- RQ3RQ.3:有影响力未来学家的语言线索如何影响社交媒体用户的前瞻性行为?
- RQ4RQ.4:‘未来现在’概念化形式在在线话语中的结构与演变是怎样的?
主要发现
- LDA模型在Twitter未来学家的推文中识别出15个广泛的主题,代表了高层次的前瞻性叙事。
- BERTopic模型发现了100个独特且语义丰富的主题,相比LDA展现出更优的粒度与可解释性。
- 绝大多数预测的未来被归类为‘未来现在’——即动态、演化且非预定的潜在未来概念化形式。
- 这些‘未来现在’型未来并非固定不变,而是代表活生生的、可塑的愿景,促使用户主动预见并积极应对。
- 有影响力未来学家的语言线索作为信号,增强了社交媒体用户预见并为多种可能未来做好准备的能力。
- 本研究证明,社交媒体数据,尤其是来自思想领袖的数据,可被有效建模,以大规模提取与分析前瞻性叙事。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。