Skip to main content
QUICK REVIEW

[论文解读] Implications of Inter-Rater Agreement on a Student Information Retrieval Evaluation

Philipp Schaer, Philipp Mayr|arXiv (Cornell University)|Oct 9, 2010
Information Retrieval and Search Behavior参考文献 9被引用 6
一句话总结

本研究使用73名信息科学专业学生的相关性判断,评估了三种基于元数据的信息检索服务——共词分析用于查询扩展、布拉德福化用于重排序、作者中心性用于文档打分。尽管评分者间一致性中等(kappa ≈ 0.4–0.6),但基于一致性阈值的数据清洗显示,三种系统检索出的文档集合彼此独立但均具相关性,凸显了在基于学生的IR评估中评分者可靠性的重要性。

ABSTRACT

This paper is about an information retrieval evaluation on three different retrieval-supporting services. All three services were designed to compensate typical problems that arise in metadata-driven Digital Libraries, which are not adequately handled by a simple tf-idf based retrieval. The services are: (1) a co-word analysis based query expansion mechanism and re-ranking via (2) Bradfordizing and (3) author centrality. The services are evaluated with relevance assessments conducted by 73 information science students. Since the students are neither information professionals nor domain experts the question of inter-rater agreement is taken into consideration. Two important implications emerge: (1) the inter-rater agreement rates were mainly fair to moderate and (2) after a data-cleaning step which erased the assessments with poor agreement rates the evaluation data shows that the three retrieval services returned disjoint but still relevant result sets.

研究动机与目标

  • 评估评分者间一致性对基于学生的信息检索评估中相关性判断的影响。
  • 评估三种基于元数据的检索服务,旨在克服tf-idf在数字图书馆中的局限性。
  • 检验不同检索机制的系统在非专家学生评判下是否检索出不同但相关的结果集合。
  • 确定移除低一致性相关性判断对整体评估结果的影响。

提出的方法

  • 使用73名信息科学专业学生作为评分者,未经事先培训进行相关性判断。
  • 应用三种检索服务:共词分析用于查询扩展,布拉德福化用于重排序,作者中心性用于文档打分。
  • 使用Fleiss’ Kappa计算评分者间一致性,以评估评分者的一致性。
  • 通过移除一致性较低的判断(kappa < 0.4)进行数据清洗,以提升数据质量。
  • 比较三种系统清洗后检索结果集合的重叠程度与相关性。

实验结果

研究问题

  • RQ1非专家学生在信息检索评估中的相关性判断一致性如何?
  • RQ2评分者间一致性对IR评估结果有效性的影响力是什么?
  • RQ3三种检索服务是否返回重叠或独立的相关文档集合?
  • RQ4移除低一致性相关性判断对评估结果有何影响?

主要发现

  • 73名学生评分者之间的评分者间一致性为中等,Fleiss’ Kappa值约为0.4至0.6。
  • 在移除一致性较差的判断(kappa < 0.4)后,剩余数据表明三种检索服务检索出的文档集合基本互不重叠。
  • 尽管一致性较低,但清洗后的数据表明,三种系统均返回了相关文档,表明其具有互补优势。
  • 结果表明,只要应用适当的清洗程序,低评分者间一致性并不必然导致评估结果无效。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。