Skip to main content
QUICK REVIEW

[论文解读] Relevance distributions across Bradford Zones: Can Bradfordizing improve search?

Philipp Mayr|arXiv (Cornell University)|May 2, 2013
Advanced Text Analysis Techniques参考文献 17被引用 11
一句话总结

本文评估了布拉德福化(Bradfordizing)这一文献计量技术,该技术将文献重新排序为核心区(第1区)和外围区(第2区和第3区),以提升信息检索的相关性。基于跨多个学科的164个标准化主题,研究结果表明,与基线及外围区相比,布拉德福化显著提高了核心区的平均精度,支持其作为A&I数据库中非文本重排序方法的应用。

ABSTRACT

The purpose of this paper is to describe the evaluation of the effectiveness of the bibliometric technique Bradfordizing in an information retrieval (IR) scenario. Bradfordizing is used to re-rank topical document sets from conventional abstracting & indexing (A&I) databases into core and more peripheral document zones. Bradfordized lists of journal articles and monographs will be tested in a controlled scenario consisting of different A&I databases from social and political sciences, economics, psychology and medical science, 164 standardized IR topics and intellectual assessments of the listed documents. Does Bradfordizing improve the ratio of relevant documents in the first third (core) compared to the second and last third (zone 2 and zone 3, respectively)? The IR tests show that relevance distributions after re-ranking improve at a significant level if documents in the core are compared with documents in the succeeding zones. After Bradfordizing of document pools, the core has a significant better average precision than zone 2, zone 3 and baseline. This paper should be seen as an argument in favour of alternative non-textual (bibliometric) re-ranking methods which can be simply applied in text-based retrieval systems and in particular in A&I databases.

研究动机与目标

  • 评估布拉德福化是否能提高排序列表前三分之一(核心区)中相关文献的比例。
  • 评估基于文献计量的重排序在不依赖文本特征的情况下,是否能提升检索精度。
  • 检验布拉德福化后,核心区在不同学术学科中的相关性分布是否优于基线和外围区。
  • 为将非文本、文献计量方法整合到基于文本的检索系统中提供实证依据。

提出的方法

  • 从社会科学、经济学、心理学和医学科学等领域的A&I数据库中,获取了164个标准化信息检索主题的文献。
  • 布拉德福化过程根据引文频率和期刊影响因子,将期刊文章和专著重新排序为三个区域,形成一个核心区(第1区)和两个外围区(第2区和第3区)。
  • 通过领域专家对文献进行知识性评估,分配相关性评分。
  • 计算每个区域的平均精度,并与核心区(第1区)、第2区、第3区及基线(未重排序)列表进行比较。
  • 应用显著性检验,比较各区域与基线之间的相关性分布差异。
  • 评估采用受控实验设置,确保各学科间主题集合和相关性判断的一致性。

实验结果

研究问题

  • RQ1与第二和第三部分(第2区和第3区)相比,布拉德福化是否显著提高了排序列表前三分之一(核心区)的平均精度?
  • RQ2基于期刊影响因子和引文频率的文献计量重排序,是否能提升基于文本的信息检索系统的有效性?
  • RQ3在多个学术学科中,布拉德福化后相关性分布的改善是否具有统计显著性?
  • RQ4核心区的相关性分布与基线(未重排序)列表相比,在平均精度方面表现如何?

主要发现

  • 布拉德福化后,所有测试学科中,核心区(第1区)的平均精度显著高于第2区和第3区。
  • 核心区的平均精度优于基线检索列表,表明布拉德福化提升了检索有效性。
  • 核心区平均精度的提升具有统计显著性,支持该方法的可靠性。
  • 各区域间的相关性分布显示出从第1区到第3区的精度明显下降,证实了布拉德福区域模型在信息检索情境下的有效性。
  • 结果表明,非文本的文献计量重排序可提升检索性能,而无需修改基于文本的索引或检索算法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。