Skip to main content
QUICK REVIEW

[论文解读] On Anomaly Ranking and Excess-Mass Curves

Nicolas Goix, Anne Sabourin|arXiv (Cornell University)|Feb 5, 2015
Anomaly Detection Techniques and Applications参考文献 12被引用 7
一句话总结

本文引入了过剩质量(Excess-Mass, EM)曲线作为多元异常值排序的新型性能准则,为最小体积曲线提供了一种功能替代方案。该方法提出了一种自适应评分函数估计器,其学习速率收敛率为 $n^{-1/2}$,即使在支撑集无界的情况下依然成立,并通过真实密度与估计密度之间的 $L^1$-距离严格控制模型偏差,从而在最小假设下实现普遍一致性。

ABSTRACT

Learning how to rank multivariate unlabeled observations depending on their degree of abnormality/novelty is a crucial problem in a wide range of applications. In practice, it generally consists in building a real valued "scoring" function on the feature space so as to quantify to which extent observations should be considered as abnormal. In the 1-d situation, measurements are generally considered as "abnormal" when they are remote from central measures such as the mean or the median. Anomaly detection then relies on tail analysis of the variable of interest. Extensions to the multivariate setting are far from straightforward and it is precisely the main purpose of this paper to introduce a novel and convenient (functional) criterion for measuring the performance of a scoring function regarding the anomaly ranking task, referred to as the Excess-Mass curve (EM curve). In addition, an adaptive algorithm for building a scoring function based on unlabeled data X1 , . . . , Xn with a nearly optimal EM is proposed and is analyzed from a statistical perspective.

研究动机与目标

  • 为无监督学习中多元异常值排序缺乏原则性性能准则的问题提供解决方案。
  • 开发一种统计一致的自适应评分函数估计器,即使在底层分布具有无界支撑集时也能表现良好。
  • 将模型偏差纳入异常值排序性能的理论分析中,从而放宽对真实密度水平集与假设类完全匹配的要求。
  • 为经验 EM 曲线优化提供 $n^{-1/2}$ 阶的学习速率上界,优于以往的 $n^{-1/4}$ 速率。

提出的方法

  • 提出过剩质量(EM)曲线作为功能性准则,通过在体积约束下最大化水平集的质量来实现,其目标与最小体积集相反。
  • 将最优 EM 曲线定义为在所有给定体积的可测集上过剩质量的上确界,与密度等高线簇一致。
  • 通过在有限阈值网格上对 EM 准则进行离散化优化,构建一组嵌套的样本密度水平集。
  • 利用一致收敛性论证和集中不等式,以高概率界定了经验 EM 曲线与真实 EM 曲线之间的偏差。
  • 通过真实密度 $f$ 与假设类 $\mathcal{F}$ 上投影 $f_F$ 之间的 $L^1$-距离来纳入模型偏差,实现偏差的统一有界控制。
  • 采用基于经验分位数的数据驱动阈值策略,确保在无参数假设下仍能保持一致性和自适应性。

实验结果

研究问题

  • RQ1能否为多元异常值排序定义一种功能性准则,使其能推广单变量尾部分析并支持一致学习?
  • RQ2在无紧支撑假设下,EM 曲线是否比现有准则(如质量-体积曲线)具有更快的收敛速率?
  • RQ3当真实密度水平集不在假设类中时,如何对异常值排序中的模型偏差进行形式化量化与控制?
  • RQ4在最小正则性假设下,是否可能实现异常值评分函数的 $n^{-1/2}$ 学习速率收敛?

主要发现

  • 所提出的 EM 曲线准则确保最优评分函数在密度的非减变换下保持不变,与异常值排序的目标一致。
  • 经验 EM 曲线优化的学习速率达到 $n^{-1/2}$,显著快于以往质量-体积曲线方法的 $n^{-1/4}$ 速率。
  • 即使在底层分布具有无界支撑集的情况下,该方法在温和正则性条件下仍保持一致,并实现 $n^{-1/2}$ 收敛速率。
  • 模型偏差由 $\|f - f_F\|_{L^1}$ 统一有界,为模型误设导致的近似误差提供了非渐近控制。
  • 理论保证通过高概率一致偏差界建立,依赖于集中不等式和 EM 曲线的导数性质。
  • 该算法生成一组嵌套的样本密度水平集,自然地实现了基于异常度评分的观测排序。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。