Skip to main content
QUICK REVIEW

[论文解读] A two-sided academic landscape: portrait of highly-cited documents in Google Scholar (1950-2013)

Alberto Martín‐Martín, Enrique Orduña‐Malea|arXiv (Cornell University)|Jul 11, 2016
scientometrics and bibliometrics research参考文献 46被引用 3
一句话总结

本研究分析了谷歌学术自1950年至2013年每年被引用次数最高的1,000篇文献,识别其类型、语言、可获取性及版本情况。研究发现,高被引文献主要为英文期刊文章和书籍,且在线PDF获取率较高,由于谷歌学术对非期刊资料的广泛覆盖,其呈现的学术图景比传统数据库更为广阔。

ABSTRACT

The main objective of this paper is to identify the set of highly-cited documents in Google Scholar and to define their core characteristics (document types, language, free availability, source providers, and number of versions), under the hypothesis that the wide coverage of this search engine may provide a different portrait about this document set respect to that offered by the traditional bibliographic databases. To do this, a query per year was carried out from 1950 to 2013 identifying the top 1,000 documents retrieved from Google Scholar and obtaining a final sample of 64,000 documents, of which 40% provided a free full-text link. The results obtained show that the average highly-cited document is a journal article or a book (62% of the top 1% most cited documents of the sample), written in English (92.5% of all documents) and available online in PDF format (86.0% of all documents). Yet, the existence of errors especially when detecting duplicates and linking cites properly must be pointed out. The fact of managing with highly cited papers, however, minimizes the effects of these limitations. Given the high presence of books, and to a lesser extend of other document types (such as proceedings or reports), the research concludes that Google Scholar data offer an original and different vision of the most influential academic documents (measured from the perspective of their citation count), a set composed not only by strictly scientific material (journal articles) but academic in its broad sense

研究动机与目标

  • 识别并描述谷歌学术中过去64年期间高被引文献的集合特征。
  • 比较谷歌学术中高被引文献的构成与传统书目数据库中的构成差异。
  • 评估非期刊文献类型(如书籍、会议论文集和报告)在学术影响力中的作用。
  • 评估全文版本(尤其是免费PDF)在谷歌学术索引中的可获取性与可访问性。
  • 考察索引限制(如重复检测和引文链接错误)对高被引文献分析可靠性的影响。

提出的方法

  • 自1950年至2013年,每年在谷歌学术中执行查询,检索每年被引用次数最高的1,000篇文献。
  • 经过去重和质量筛选后,最终数据集包含64,000篇唯一文献。
  • 提取了文献特征,包括类型(如期刊文章、书籍、会议论文)、语言、免费全文链接的可获取性、来源提供方及版本数量。
  • 应用统计分析方法,评估样本中各类文献类型、语言分布及数字可获取性的分布情况。
  • 通过评估重复检测和引文链接准确性的方法,评估谷歌学术引文索引的可靠性。
  • 将研究发现与传统书目数据库进行对比,突出高被引文献在呈现上的差异。

实验结果

研究问题

  • RQ1与传统数据库相比,谷歌学术中高被引文献的主导文献类型是什么?
  • RQ2谷歌学术中高被引文献的语言和格式分布(如英文、PDF)如何?
  • RQ3在谷歌学术中,高被引文献的免费全文版本普遍程度如何?
  • RQ4在以引文次数衡量的学术影响力中,非期刊文献(如书籍和报告)的贡献程度如何?
  • RQ5谷歌学术中的索引错误(如重复检测、引文链接错误)在多大程度上影响了识别高被引文献的可靠性?

主要发现

  • 样本中被引次数最高的1%的文献中,62%为期刊文章或书籍,表明其在学术影响力中的主导地位。
  • 所有高被引文献中92.5%以英文发表,凸显高影响力学术成果中的语言霸权现象。
  • 86.0%的文献以PDF格式提供,反映出谷歌学术在数字可访问性方面的强大支持。
  • 样本中64,000篇文献中有40%提供了免费全文链接,表明显著的开放获取可及性。
  • 与传统引文数据库相比,书籍及其他非期刊文献类型(如会议录、技术报告)在样本中被显著高估。
  • 尽管存在重复检测和引文链接的索引错误,但文献的高被引次数在整体上缓解了这些局限性对研究结果的影响。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。