Skip to main content
QUICK REVIEW

[论文解读] The Rise of Open Science: Tracking the Evolution and Perceived Value of Data and Methods Link-Sharing Practices

Hancheng Cao, Jesse Dodge|arXiv (Cornell University)|Oct 4, 2023
Scientific Computing and Data Management被引用 5
一句话总结

本研究利用神经文本分类模型识别论文中的URL类型,分析了110万篇来自arXiv的计算机科学、物理学和数学论文中数据与方法链接共享的演变及其感知价值。研究发现,链接共享的采用率随时间持续上升,链接复用率(尤其是GitHub平台)不断增长,且拥有活跃数据/方法链接的论文引用率更高,表明开放科学实践正获得日益提升的机构与学术价值。

ABSTRACT

In recent years, funding agencies and journals increasingly advocate for open science practices (e.g. data and method sharing) to improve the transparency, access, and reproducibility of science. However, quantifying these practices at scale has proven difficult. In this work, we leverage a large-scale dataset of 1.1M papers from arXiv that are representative of the fields of physics, math, and computer science to analyze the adoption of data and method link-sharing practices over time and their impact on article reception. To identify links to data and methods, we train a neural text classification model to automatically classify URL types based on contextual mentions in papers. We find evidence that the practice of link-sharing to methods and data is spreading as more papers include such URLs over time. Reproducibility efforts may also be spreading because the same links are being increasingly reused across papers (especially in computer science); and these links are increasingly concentrated within fewer web domains (e.g. Github) over time. Lastly, articles that share data and method links receive increased recognition in terms of citation count, with a stronger effect when the shared links are active (rather than defunct). Together, these findings demonstrate the increased spread and perceived value of data and method sharing practices in open science.

研究动机与目标

  • 追踪物理学、数学和计算机科学领域内科学论文中数据与方法链接共享实践的演变过程。
  • 量化数据与方法链接随时间推移的采用率与复用模式。
  • 通过分析链接共享与引用影响力的关联,评估链接共享的感知价值。
  • 识别网络域名(如GitHub)在集中链接共享活动中的作用。

提出的方法

  • 训练一个神经文本分类模型,基于上下文提及(如“data”、“code”、“method”)自动分类论文中的URL类型。
  • 将该模型应用于arXiv平台中涵盖计算机科学、物理学和数学领域的110万篇论文的大规模数据集。
  • 追踪链接共享频率、跨论文复用情况以及域名集中度(如GitHub、Zenodo)的时间趋势。
  • 测量共享数据/方法链接的论文的引用影响力,并对比活跃链接与失效链接的差异。
  • 使用统计分析方法,关联链接共享实践与论文接受度,包括引用次数。

实验结果

研究问题

  • RQ1在arXiv的科学论文中,数据与方法链接共享率随时间如何演变?
  • RQ2共享的数据与方法链接在不同论文间的复用程度如何?这种复用随时间如何变化?
  • RQ3哪些网络域名最常用于托管共享的数据与方法?这种集中趋势如何演变?
  • RQ4链接共享与引用影响力之间存在何种关系,特别是对比活跃链接与失效链接时?
  • RQ5在计算机科学、物理学和数学三个领域中,链接共享实践有何差异?

主要发现

  • 共享数据与方法链接的论文比例随时间显著增加,尤其在计算机科学领域更为明显。
  • 共享链接在不同论文间的复用日益频繁,表明可重现研究实践的采用率持续上升。
  • 链接共享正更加集中于少数网络域名,GitHub已成为方法与数据共享的主导平台。
  • 拥有活跃数据与方法链接的论文获得的引用次数显著高于链接失效或缺失的论文。
  • 活跃链接对引用数的正向影响更强,表明链接质量与持久性对学术价值的感知至关重要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。