[论文解读] Measuring agreement among several raters classifying subjects into one or more (hierarchical) categories: A generalization of Fleiss' kappa
本文提出了一种广义的 Fleiss' kappa 统计量,用于在多个评分者将受试者分类到一个或多个名义类别(包括分层和加权类别)时衡量评分者间的一致性。该方法通过同时考虑选择与非选择的一致性,扩展了 Fleiss' kappa,确保在评分者分配多个类别时仍能保持校正偶然一致性的可靠性,并在单类别分类情况下与 Fleiss' kappa 保持一致。
Cohen's and Fleiss' kappa are well-known measures of inter-rater agreement, but they restrict each rater to selecting only one category per subject. This limitation is consequential in contexts where subjects may belong to multiple categories, such as psychiatric diagnoses involving multiple disorders or classifying interview snippets into multiple codes of a codebook. We propose a generalized version of Fleiss' kappa, which accommodates multiple raters assigning subjects to one or more nominal categories. Our proposed $κ$ statistic can incorporate category weights based on their importance and account for hierarchical category structures, such as primary disorders with sub-disorders. The new $κ$ statistic can also manage missing data and variations in the number of raters per subject or category. We review existing methods that allow for multiple category assignments and detail the derivation of our measure, proving its equivalence to Fleiss' kappa when raters select a single category per subject. The paper discusses the assumptions, premises, and potential paradoxes of the new measure, as well as the range of possible values and guidelines for interpretation. The measure was developed to investigate the reliability of a new mathematics assessment method, of which an example is elaborated. The paper concludes with a worked-out example of psychiatrists diagnosing patients with multiple disorders. All calculations are provided as R script and an Excel sheet to facilitate access to the new $κ$ tatistic.
研究动机与目标
- 为解决现有评分者间一致性度量(如 Cohen’s 和 Fleiss’ kappa)的局限性,这些度量要求评分者为每个受试者恰好分配一个类别。
- 开发一种校正偶然一致性的可靠性度量,允许分配多个类别,同时保持可解释性和统计严谨性。
- 在一致性计算中纳入类别之间的分层依赖关系(例如,主要障碍与子障碍)以及类别的加权重要性。
- 确保当每个受试者仅分配一个类别时,该度量与 Fleiss’ kappa 保持一致。
- 在精神病学诊断和教育评估等研究场景中实现实际应用,这些场景中多个诊断或标准的分配很常见。
提出的方法
- 所提出的 kappa 统计量通过计算评分者之间成对一致性的平均值来计算观察到的一致性 $ P_o $,并针对多个类别分配进行调整。
- 基于所有评分者和受试者的类别分配边缘概率,定义了由机会导致的一致性 $ P_e $。
- 采用对称的一致性定义:若两个评分者均选择或均未选择同一类别,则视为一致,无论类别数量如何。
- 通过允许 $ x_{ic} $ 表示将类别 $ c $ 分配给受试者 $ i $ 的评分者数量,将 Fleiss 的公式推广到 $ \sum_c x_{ic} > J $ 的情况。
- 该方法支持加权类别(例如,按临床重要性加权)和分层结构(例如,子类别仅在父类别被选择时才可用)。
- 该方法可处理缺失数据和每个受试者或类别的评分者数量变化,增强了在真实世界数据中的稳健性。

实验结果
研究问题
- RQ1当评分者为每个受试者分配多个类别而非仅一个时,如何衡量评分者间的一致性?
- RQ2当类别具有分层或加权结构时,适当的校正偶然一致性的相关系数是什么?
- RQ3与比例重叠或校正偶然一致性的相关性等现有方法相比,所提出的度量在一致性和可解释性方面表现如何?
- RQ4当每个受试者仅分配一个类别时,所提出的统计量是否能与 Fleiss’ kappa 保持等价?
- RQ5在可靠性分析中忽略非选择的一致性(即两个评分者均未选择某一类别)会产生什么实际影响?
主要发现
- 当每个受试者由所有评分者恰好分配一个类别时,所提出的 kappa 统计量退化为 Fleiss’ kappa,确保与既定方法的一致性。
- 该度量同时考虑了选择与非选择的一致性,避免了忽略非选择而引入的偏差。
- 该方法支持分层类别结构,例如主要精神障碍及其相关子障碍,通过建模类别之间的依赖关系实现。
- 可整合加权类别,使研究人员能根据上下文为不同类别分配不同的重要性水平。
- 该方法可处理缺失数据和每个受试者或类别的评分者数量变化,显著提升了实际应用价值。
- 作者计划发布一个 R 包以促进其应用,并设想通过模拟研究证明其均方根误差低于其他替代方法。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。