Skip to main content
QUICK REVIEW

[论文解读] Developing indicators on Open Access by combining evidence from diverse data sources

Thed N. van Leeuwen, Ingeborg Meijer|arXiv (Cornell University)|Feb 8, 2018
scientometrics and bibliometrics research被引用 8
一句话总结

本文提出了一种系统化方法,通过整合多种数据源(如出版商元数据、机构知识库和DOI)对出版物进行开放获取(OA)标签标注,从而克服Web of Science等书目数据库中OA标签不完整的问题。该方法显著提高了OA检测的准确性,并支持对研究产出中OA趋势的可靠追踪,有助于科学政策制定与管理。

ABSTRACT

In the last couple of years, the role of Open Access (OA) publishing has become central in science management and research policy. In the UK and the Netherlands, national OA mandates require the scientific community to seriously consider publishing research outputs in OA forms. At the same time, other elements of Open Science are becoming also part of the debate, thus including not only publishing research outputs but also other related aspects of the chain of scientific knowledge production such as open peer review and open data. From a research management point of view, it is important to keep track of the progress made in the OA publishing debate. Until now, this has been quite problematic, given the fact that OA as a topic is hard to grasp by bibliometric methods, as most databases supporting bibliometric data lack exhaustive and accurate open access labelling of scientific publications. In this study, we present a methodology that systematically creates OA labels for large sets of publications processed in the Web of Science database. The methodology is based on the combination of diverse data sources that provide evidence of publications being OA

研究动机与目标

  • 解决Web of Science等主要书目数据库中开放获取(OA)标签不一致、不完整的问题。
  • 开发一种系统化、可扩展的方法,利用多种证据来源在大规模出版物集合中识别OA出版物。
  • 通过实现对OA出版趋势及国家OA指令合规性的准确追踪,支持科学政策与研究管理。
  • 超越传统文献计量指标,整合更广泛的开放科学维度,如开放数据与开放同行评审。
  • 为跨机构和国家生成可靠OA指标提供可复现的框架。

提出的方法

  • 汇集来自多种数据源的证据,包括出版商元数据、机构知识库(如PubMed Central、机构知识库)以及带有OA状态的DOI。
  • 使用基于规则的分类系统,根据所收集数据源中是否存在OA指标来分配OA状态。
  • 应用分层决策流程:若出版物出现在可信的OA知识库中,或具有明确的OA许可协议,则标记为OA。
  • 通过已知的OA数据集验证该方法,以确保标签过程的准确性和可靠性。
  • 使用该多源证据方法处理来自Web of Science数据库的大规模出版物集合。
  • 实施评分机制,根据证据来源的数量和可靠性评估OA标签的置信度。

实验结果

研究问题

  • RQ1当书目数据库缺乏全面OA标签时,如何可靠地确定大规模出版物集合的开放获取状态?
  • RQ2哪些数据源组合能提供最准确且可扩展的OA出版物识别方法?
  • RQ3与单源方法相比,多源证据方法在多大程度上提升了OA检测的准确性?
  • RQ4该方法如何支持国家和机构对OA合规性及开放科学采纳情况的监测?
  • RQ5该方法是否具有普适性,可否应用于不同学科和出版类型?

主要发现

  • 与依赖单一来源元数据相比,多源证据方法显著提高了OA检测的准确性,尤其在OA标签不完整的数据库中表现更优。
  • 通过结合机构知识库、出版商网站和DOI的信号,该方法能以高精度识别OA出版物。
  • 该方法实现了大规模出版物集合的一致且可复现的标签标注,支持对OA趋势的纵向监测。
  • 研究表明,整合多样化数据源可减少OA识别中的漏报,尤其对混合型或延迟OA期刊中的出版物效果显著。
  • 该方法支持生成可靠的OA指标,用于政策评估和机构报告,如英国和荷兰OA指令的合规性监测。
  • 置信度评分系统可对OA标签质量进行评估,使研究人员能在分析中优先使用高置信度的OA记录。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。