Skip to main content
QUICK REVIEW

[论文解读] An Investigation into Outlier Elimination and Calculation Methods in the Determination of Reference Intervals using Serum Immunoglobulin A as a Model Data Collection

Aidan Zellner, Alice Richardson|arXiv (Cornell University)|Jul 23, 2019
Clinical Laboratory Practices and Quality Control参考文献 24被引用 7
一句话总结

本研究基于超过32,000项血清免疫球蛋白A(IgA)检测数据,调查了异常值剔除方法与参考区间计算方法。研究发现,异常值剔除方法——尤其是Tukey方法——对参考区间确定的影响远大于计算方法的影响。Tukey方法相较于分块法(Dixon/Reed)剔除的异常值显著更多,当应用该方法后,不同计算技术之间的差异极小。

ABSTRACT

Background: Reference intervals are essential to interpret diagnostic tests, but their determination has become controversial. Methods: In this paper parametric, non-parametric and robust reference intervals with Tukey and block elimination are calculated from a dataset of over 32,000 serum immunoglobulin A (IgA) measurements. Results: The outlier elimination method was significantly more determinative of the reference intervals than the calculation method. The Tukey elimination procedure consistently eliminated significantly more values than the block method of Dixon and Reed across all age ranges. If Tukey elimination was applied, variation between reference intervals produced by the different calculation methods was minimal. Block elimination rarely eliminated values. The non-parametric reference intervals were more sensitive to outliers, which in the IgA context, led to higher and wider reference intervals for the older age groups. There were only minimal differences between robust and parametric reference intervals. Conclusions: This suggests that Tukey elimination should be preferred over the block D/R method for datasets similar to the one used in this study. These are predominantly new observations, as previous literature has focused on the calculation technique and not discussed outlier elimination. This suggests the robust method is not advantageous over the parametric method and therefore due to its complexity is not particularly useful, contrary to CLSI Guidelines.

研究动机与目标

  • 评估不同异常值剔除技术对临床实验室数据中参考区间确定的影响。
  • 比较参数法、非参数法与稳健统计方法在计算参考区间中的表现。
  • 评估在真实世界临床数据集中,稳健方法是否相较于参数方法具有优势。
  • 通过检验其实际影响,挑战CLSI指南中关于稳健方法必要性的既定假设。
  • 确定异常值剔除或计算方法对最终参考区间值的影响哪个更大。

提出的方法

  • 将参数法、非参数法与稳健参考区间计算方法应用于包含32,000多项血清IgA测量值的数据集。
  • 采用Tukey的内边界与外边界方法进行异常值检测与剔除。
  • 应用Dixon与Reed的分块法进行异常值检测,并与Tukey方法的性能进行比较。
  • 在多个年龄组中计算参考区间,以评估对异常值处理的敏感性。
  • 使用各方法的标准统计公式计算参考区间:参数法(均值 ± 1.96标准差)、非参数法(2.5百分位数与97.5百分位数)、稳健法(中位数与四分位距,含修剪处理)。
  • 比较不同方法与异常值剔除技术下生成的参考区间,以评估其变异性与一致性。

实验结果

研究问题

  • RQ1在大型IgA数据集中,不同的异常值剔除方法(Tukey法 vs. 分块法)如何影响最终的参考区间值?
  • RQ2当与异常值剔除结合时,哪种计算方法——参数法、非参数法或稳健法——能产生最稳定可靠的参考区间?
  • RQ3异常值剔除方法的选择对参考区间确定的影响是否大于计算方法的选择?
  • RQ4当存在异常值时,非参数参考区间与参数及稳健区间在多大程度上存在差异?
  • RQ5正如CLSI指南所声称的那样,稳健方法在真实世界临床数据中是否相较于参数方法具有实际优势?

主要发现

  • 异常值剔除方法对参考区间值的影响显著大于计算方法的选择。
  • 在所有年龄组中,Tukey异常值剔除程序剔除的数值远多于分块法(Dixon/Reed)的剔除数量。
  • 当应用Tukey剔除后,由参数法、非参数法与稳健法生成的参考区间差异极小。
  • Dixon与Reed的分块法极少剔除任何数值,导致区间估计几乎无变化。
  • 非参数参考区间对异常值更敏感,尤其在老年组中产生更高且更宽的区间。
  • 稳健法与参数法生成的参考区间仅存在微小差异,表明在此数据集中,稳健方法的额外复杂性并无实际优势。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。