Skip to main content
QUICK REVIEW

[论文解读] A Two-stage Online Monitoring Procedure for High-Dimensional Data Streams

Jun Li|arXiv (Cornell University)|Dec 14, 2017
Advanced Statistical Process Monitoring参考文献 9被引用 3
一句话总结

本文提出了一种针对高维数据流的两阶段在线监控程序,可同时将受控状态下的平均运行长度(IC-ARL)和第一类错误率(假发现率,FDR)控制在用户指定的水平。通过将误报控制与检测效能分离,该方法克服了现有单阶段方法的局限性,这些方法要么缺乏全局FDR控制,要么在IC-ARL与检测灵敏度之间强制权衡。

ABSTRACT

Advanced computing and data acquisition technologies have made possible the collection of high-dimensional data streams in many fields. Efficient online monitoring tools which can correctly identify any abnormal data stream for such data are highly sought after. However, most of the existing monitoring procedures directly apply the false discover rate (FDR) controlling procedure to the data at each time point, and the FDR at each time point (the point-wise FDR) is either specified by users or determined by the in-control (IC) average run length (ARL). If the point-wise FDR is specified by users, the resulting procedure lacks control of the global FDR and keeps users in the dark in terms of the IC-ARL. If the point-wise FDR is determined by the IC-ARL, the resulting procedure does not give users the flexibility to choose the number of false alarms (Type-I errors) they can tolerate when identifying abnormal data streams, which often makes the procedure too conservative. To address those limitations, we propose a two-stage monitoring procedure that can control both the IC-ARL and Type-I errors at the levels specified by users. As a result, the proposed procedure allows users to choose not only how often they expect any false alarms when all data streams are IC, but also how many false alarms they can tolerate when identifying abnormal data streams. With this extra flexibility, our proposed two-stage monitoring procedure is shown in the simulation study and real data analysis to outperform the exiting methods.

研究动机与目标

  • 解决现有针对高维数据流的在线监控程序中缺乏全局假发现率(FDR)控制的问题。
  • 解决当前方法将逐点FDR与受控状态下的平均运行长度(IC-ARL)绑定所带来的不灵活性,从而限制了用户对误报容忍度的控制。
  • 开发一种监控方案,使用户能够在识别失控数据流时独立指定期望的IC-ARL和第一类错误率水平。
  • 通过引入两阶段决策框架,提升高维流数据中的检测效能与误报控制。

提出的方法

  • 提出一种两阶段监控策略:第一阶段使用CUSUM统计量为每个数据流计算p值,第二阶段对这些p值应用FDR控制程序。
  • 提出一种新颖的两阶段决策规则,将受控状态下的平均运行长度(IC-ARL)控制与第一类错误率(FDR)控制分离,从而实现用户对两者的独立指定。
  • 采用分层检验框架,通过调整第二阶段的阈值来控制全局FDR,同时通过校准第一阶段的控制限来维持IC-ARL。
  • 采用基于似然比的方法证明,该两阶段程序在似然比上具有单调性,从而在原假设与备择假设下均能确保有效的FDR控制。
  • 通过原假设与备择假设下的概率密度函数,推导出p值与检验统计量之间的关系,从而将监控问题转化为p值空间中的问题。
  • 应用p值变换,利用检验统计量之间的依赖结构,确保FDR在整个时间窗口内得到全局控制,而不仅是在每个时间点的逐点控制。

实验结果

研究问题

  • RQ1在文献中普遍假设的前提下,仅在每个时间点控制点态FDR是否能保证在整个时间窗口内实现全局FDR控制?
  • RQ2是否存在一种监控程序,能够同时将受控状态下的平均运行长度(IC-ARL)和第一类错误率(FDR)控制在用户指定的水平?
  • RQ3在高维流数据中,该两阶段程序相较于单阶段方法在检测效能与误报控制方面表现如何?
  • RQ4在实际应用中,用户指定的IC-ARL与FDR水平对监控程序性能有何影响?

主要发现

  • 本文证明,仅控制点态FDR并不能保证全局FDR的控制,从而挑战了文献中广泛持有的假设。
  • 所提出的两阶段程序成功地在用户指定的水平上控制了IC-ARL与全局FDR,相较于现有单阶段方法具有更高的灵活性。
  • 模拟研究显示,当数据流数量较大时,该两阶段方法在检测速度与误报控制方面均优于现有方法。
  • 该方法在保持期望IC-ARL的同时,允许用户选择其偏好的第一类错误容忍度,避免了通过IC-ARL固定点态FDR而产生的保守性权衡。
  • 对网络流量与健康监测数据的真实数据分析证实,该方法在识别失控流方面表现更优,且误报更少。
  • 理论证明表明,检验统计量的似然比具有单调性,从而确保了在两阶段决策规则下FDR控制的有效性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。