[论文解读] Evaluation of the citation matching algorithms of CWTS and iFQ in comparison to Web of Science
本研究利用人工验证的语料库,评估了CWTS和iFQ的引文匹配算法与Web of Science(WoS)的对比表现,衡量各系统在纠正引文信息错误方面的有效性。CWTS的算法表现最佳(F1:96.41%),iFQ紧随其后,而当引文包含错误时,WoS表现出明显不足,凸显了在书目计量研究中采用稳健匹配方法的重要性。
The results of bibliometric studies provided by bibliometric research groups, e.g. the Centre for Science and Technology Studies (CWTS) and the Institute for Research Information and Quality Assurance (iFQ), are often used in the process of research assessment. Their databases use Web of Science (WoS) citation data, which they match according to their own matching algorithms - in the case of CWTS for standard usage in their studies and in the case of iFQ on an experimental basis. Since the problem of non-matched citations in WoS persists because of inaccuracies in the references or inaccuracies introduced in the data extraction process, it is important to ascertain how well these inaccuracies are rectified in these citation matching algorithms. This paper evaluates the algorithms of CWTS and iFQ in comparison to WoS in a quantitative and a qualitative analysis. The analysis builds upon the methodology and the manually verified corpus of a previous study. The algorithm of CWTS performs best, closely followed by that of iFQ. The WoS algorithm still performs quite well (F1 score: 96.41 percent), but shows deficits in matching references containing inaccuracies. An additional problem is posed by incorrectly provided cited reference information in source articles by WoS.
研究动机与目标
- 评估CWTS和iFQ所使用的引文匹配算法在纠正来自Web of Science的引文信息错误方面的表现。
- 识别Web of Science自身引文匹配流程中的系统性缺陷,特别是在源参考文献包含错误时的表现。
- 使用人工验证的参考文献语料库,对比CWTS和iFQ算法与WoS在精确率、召回率和F1得分方面的有效性。
- 评估各系统在处理数据提取或参考文献录入过程中引入的真实世界引文错误方面的表现。
- 为基于这些系统生成的书目计量数据在研究评估背景下的可靠性提供实证证据。
提出的方法
- 利用先前研究中的人工验证引文对语料库作为评估的基准真实数据。
- 将CWTS、iFQ和Web of Science的引文匹配算法应用于同一组参考文献,以比较匹配结果。
- 为每个系统的性能计算标准信息检索指标——精确率、召回率和F1得分。
- 同时开展定量分析(使用F1得分)和对匹配错误及模式的定性评估。
- 重点关注引文信息存在错误(如作者、题名、DOI缺失或错误)的情况,以评估其错误纠正能力。
- 采用与先前研究相同的评估框架,以确保结果的一致性和可比性。
实验结果
研究问题
- RQ1在人工验证的引文语料库上,CWTS和iFQ的引文匹配算法与Web of Science在F1得分方面表现如何比较?
- RQ2CWTS和iFQ在多大程度上能够纠正Web of Science参考数据中存在的引文错误?
- RQ3各系统算法最常误分类或遗漏的引文错误类型有哪些?
- RQ4当面对错误或不完整的引文信息时,Web of Science自身的匹配算法是否表现可靠?
- RQ5CWTS和iFQ在处理复杂或模糊的引文格式时,性能特征有何不同?
主要发现
- CWTS引文匹配算法取得了96.41%的最高F1得分,优于iFQ和Web of Science。
- iFQ算法表现紧随CWTS之后,表明其在纠正引文错误方面具备强大的能力。
- Web of Science自身的算法虽取得较高但非完美的F1得分(96.41%),但在引文信息存在错误时表现出显著缺陷。
- Web of Science的一个关键局限是无法有效纠正原始文献中引入的错误或数据提取过程中的错误。
- 本研究证实,引文匹配算法对书目计量准确性至关重要,尤其是在源数据存在错误时。
- 结果表明,在研究评估背景下,CWTS和iFQ在处理存在缺陷的引文数据方面,比Web of Science提供了更稳健的解决方案。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。