[论文解读] The Pushshift Reddit Dataset
本文介绍了 Pushshift Reddit 数据集,这是一个自2015年起持续更新、公开可访问的大型数据集,包含2005年至2019年间超过56亿条Reddit评论和6.51亿条帖子。该数据集为研究人员提供了实时API访问、全文搜索、聚合工具以及用于交互式分析的Slackbot,显著减少了数据收集的时间与技术负担,使后API时代的大规模社交媒体研究具备可重现性。
Social media data has become crucial to the advancement of scientific understanding. However, even though it has become ubiquitous, just collecting large-scale social media data involves a high degree of engineering skill set and computational resources. In fact, research is often times gated by data engineering problems that must be overcome before analysis can proceed. This has resulted recognition of datasets as meaningful research contributions in and of themselves. Reddit, the so called "front page of the Internet," in particular has been the subject of numerous scientific studies. Although Reddit is relatively open to data acquisition compared to social media platforms like Facebook and Twitter, the technical barriers to acquisition still remain. Thus, Reddit's millions of subreddits, hundreds of millions of users, and hundreds of billions of comments are at the same time relatively accessible, but time consuming to collect and analyze systematically. In this paper, we present the Pushshift Reddit dataset. Pushshift is a social media data collection, analysis, and archiving platform that since 2015 has collected Reddit data and made it available to researchers. Pushshift's Reddit dataset is updated in real-time, and includes historical data back to Reddit's inception. In addition to monthly dumps, Pushshift provides computational tools to aid in searching, aggregating, and performing exploratory analysis on the entirety of the dataset. The Pushshift Reddit dataset makes it possible for social media researchers to reduce time spent in the data collection, cleaning, and storage phases of their projects.
研究动机与目标
- 解决研究人员因平台API限制和隐私丑闻而日益难以获取大规模社交媒体数据的问题。
- 降低社交媒体研究中与数据收集、清洗和存储相关的耗时与技术开销。
- 通过提供具备增强功能的可信第三方数据存储库,为官方平台API提供可持续、开放且可扩展的替代方案。
- 通过广泛提供历史Reddit数据并配备高级查询与分析工具,支持计算社会科学中的可重现性与可及性。
- 使跨学科研究人员能够研究仇恨言论、极端化、心理健康和虚假信息等现象,而无需依赖受限或已弃用的平台API。
提出的方法
- Pushshift自2015年起实时收集Reddit数据,保留了自Reddit创立以来的完整帖子与评论档案。
- 数据集通过每月发布的数据包形式提供,存储于 https://files.pushshift.io/reddit/,确保长期可访问性。
- 实时可搜索API使研究人员无需下载大文件即可查询整个数据集,降低存储与计算需求。
- API支持对评论和帖子的全文搜索、汇总统计的聚合端点,且查询限制(每请求最多500个对象)是Reddit原生API的五倍。
- 通过与Slackbot集成,研究人员可直接在Slack中执行实时数据查询并生成可视化结果,提升协作与探索性分析效率。
- 平台通过提供抽象低级数据工程任务的工具,支持可重现研究,使研究人员能够专注于分析与假设检验。
实验结果
研究问题
- RQ1研究人员如何在不依赖受速率限制或已弃用的官方API的情况下,高效访问并分析大规模历史Reddit数据?
- RQ2支持可持续、开放且可扩展的社会媒体数据共享以供学术研究,所需的必要技术与基础设施组件是什么?
- RQ3第三方数据平台在多大程度上能够降低计算资源有限或技术能力不足的研究人员的入门门槛?
- RQ4提供实时API访问、全文搜索与交互式工具,如何提升计算社会科学研究的可重现性与效率?
- RQ5像Pushshift这样的社区驱动数据平台,在支持仇恨言论、心理健康与虚假信息等敏感议题的跨学科研究方面,可发挥何种作用?
主要发现
- Pushshift Reddit数据集包含自2005年Reddit创立至2019年期间收集的6.51亿条帖子与56亿条评论,构成一份全面的历史档案。
- 该数据集通过实时可搜索API提供,支持全文搜索与聚合功能,使研究人员无需本地存储即可高效查询。
- Pushshift的API支持的查询限制是Reddit原生API的五倍,使研究人员能以更少请求获取更大数据批次。
- 该平台提供Slackbot,使研究人员能够实时执行查询并生成可视化结果,提升协作分析能力。
- 已有超过100篇经同行评审的学术论文在不同学科中使用了Pushshift Reddit数据集,验证了其效用与可靠性。
- 该数据集及相关工具显著减少了数据收集、清洗与存储的时间与技术投入,使研究人员能够专注于分析与假设检验。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。