Skip to main content
QUICK REVIEW

[论文解读] Nonparametric Detection of Anomalous Data via Kernel Mean Embedding.

Shaofeng Zou, Yingbin Liang|arXiv (Cornell University)|Apr 25, 2014
Statistical Methods and Inference参考文献 21被引用 8
一句话总结

本文提出了一种基于再生核希尔伯特空间(RKHS)中核均值嵌入的非参数异常检测方法,利用最大均值差异(MMD)检测n个序列中的异常序列。当s已知时,证明每序列m = O(log n)个样本足以实现一致检测;当s未知时,需满足m > O(log n)。该方法具有多项式时间复杂度,并在传统方法和基于核的方法上展现出优异的实验性能。

ABSTRACT

An anomaly detection problem is investigated, in which there are totally n sequences with s anomalous sequences to be detected. Each normal sequence contains m independent and identically distributed (i.i.d.) samples drawn from a distribution p, whereas each anomalous sequence contains m i.i.d. samples drawn from a distribution q that is distinct from p. The distributions p and q are assumed to be unknown a priori. Two scenarios, respectively with and without a reference sequence generated by p, are studied. Distribution-free tests are constructed using maximum mean discrepancy (MMD) as the metric, which is based on mean embeddings of distributions into a reproducing kernel Hilbert space (RKHS). For both scenarios, it is shown that as the number n of sequences goes to infinity, if the value of s is known, then the number m of samples in each sequence should be at the order O(log n) or larger in order for the developed tests to consistently detect s anomalous sequences. If the value of s is unknown, then m should be at the order strictly larger than O(log n). Computational complexity of all developed tests is shown to be polynomial. Numerical results demonstrate that our tests outperform (or perform as well as) the tests based on other competitive traditional statistical approaches and kernel-based approaches under various cases. Consistency of the proposed test is also demonstrated on a real data set.

研究动机与目标

  • 解决在底层分布p和q未知时,从n个序列中检测出s个异常序列的挑战。
  • 基于最大均值差异(MMD)开发无需分布假设的检验方法,适用于存在来自分布p的参考序列与不存在参考序列两种情形。
  • 建立在已知和未知s条件下实现一致检测的理论样本量要求。
  • 通过保持所有所提检验的多项式时间复杂度,确保计算效率。
  • 在合成数据和真实数据上,通过与传统和基于核的统计方法对比,验证该方法的性能。

提出的方法

  • 使用核均值嵌入将概率分布p和q表示为再生核希尔伯特空间(RKHS)中的元素,从而实现非参数比较。
  • 采用最大均值差异(MMD)作为正常序列与异常序列的均值嵌入之间的统计距离度量。
  • 基于MMD构建无需分布假设的假设检验,以检测与正常分布p的偏离。
  • 设计两种检验变体:一种使用来自p的参考序列,另一种不使用,相应调整检验统计量。
  • 推导每序列样本数m的理论条件,以确保当n → ∞时检测的一致性。
  • 分析计算复杂度,并证明所有所提检验的复杂度在n和m上均保持多项式。

实验结果

研究问题

  • RQ1当s已知时,在p和q未知的条件下,每序列检测s个异常序列所需的最少样本数m是多少?
  • RQ2当s未知时,所需样本数m如何变化?可提供何种理论保证?
  • RQ3MMD-based检验是否可在不假设p和q参数形式的前提下实现一致异常检测?
  • RQ4与传统和基于核的异常检测方法相比,该方法在性能和复杂度上表现如何?
  • RQ5该方法在具有未知底层分布的真实数据上是否保持一致性和有效性?

主要发现

  • 当s已知时,n → ∞下每序列m = O(log n)个样本足以实现s个异常序列的一致检测。
  • 当s未知时,必须满足m > O(log n)才能确保一致检测。
  • 所有所提检验均具有多项式时间复杂度,使其在n较大时仍具可扩展性。
  • 数值实验表明,所提MMD-based检验在各种设置下性能优于或匹配传统和基于核的统计方法。
  • 该方法在真实世界数据集上表现出一致性,证实其实际可行性与理论合理性。
  • 理论框架成功处理了存在与不存在参考序列两种情形,提供了在分布不确定性下的鲁棒检测。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。