[论文解读] Learning in the Presence of Corruption
本文提出了一种基于统计决策理论的一般性框架,用于在标签噪声条件下分析监督学习,推导出基于遍历系数的泛化风险上界与下界,并提出一种广义的无偏估计方法。该框架在数据被污染的情况下仍能保持快速学习速率,特别证明了在对称标签噪声下,可分离数据的0-1损失仍能以快速率学习。
In supervised learning one wishes to identify a pattern present in a joint distribution $P$, of instances, label pairs, by providing a function $f$ from instances to labels that has low risk $\mathbb{E}_{P}\ell(y,f(x))$. To do so, the learner is given access to $n$ iid samples drawn from $P$. In many real world problems clean samples are not available. Rather, the learner is given access to samples from a corrupted distribution $ ilde{P}$ from which to learn, while the goal of predicting the clean pattern remains. There are many different types of corruption one can consider, and as of yet there is no general means to compare the relative ease of learning under these different corruption processes. In this paper we develop a general framework for tackling such problems as well as introducing upper and lower bounds on the risk for learning in the presence of corruption. Our ultimate goal is to be able to make informed economic decisions in regards to the acquisition of data sets. For a certain subclass of corruption processes (those that are \emph{reconstructible}) we achieve this goal in a particular sense. Our lower bounds are in terms of the coefficient of ergodicity, a simple to calculate property of stochastic matrices. Our upper bounds proceed via a generalization of the method of unbiased estimators appearing in recent work of Natarajan et al and implicit in the earlier work of Kearns.
研究动机与目标
- 开发一个一般性的理论框架,用于比较不同类型污染数据的学习难度。
- 在训练数据被污染的情况下,提供泛化风险的上下界,以实现对数据质量的定量比较。
- 通过评估不同污染数据类型的效用,支持关于数据获取的经济决策。
- 刻画在何种条件下快速学习速率在污染数据下仍能保持,特别是针对0-1损失和标签噪声。
- 形式化描述在学习效率方面,污染数据可被视为与干净数据等价的条件。
提出的方法
- 将污染过程建模为从干净观测空间到污染观测空间的马尔可夫核,使用有限集合和随机矩阵。
- 应用随机矩阵的遍历系数作为污染严重程度的关键度量,以推导风险的下界。
- 将先前工作中提出的无偏估计方法推广,以构建对污染具有鲁棒性的学习算法。
- 引入损失与污染核之间的η-相容性概念,以关联干净数据与污染数据的风险边界。
- 利用组合性引理分析多个污染过程的组合(例如,标签噪声与部分标签)。
- 通过广义无偏估计构造方法推导上界,通过遍历系数推导下界,两者均基于Bernstein条件。
实验结果
研究问题
- RQ1在何种条件下,从污染数据中学习时仍能保持快速学习速率?
- RQ2如何对不同类型污染数据(如标签噪声、部分标签)在学习效用方面进行定量比较?
- RQ3污染过程(作为马尔可夫核)与最终泛化风险之间存在何种关系?
- RQ4Bernstein条件在污染下是否仍能保持?若能,其条件是什么?
- RQ5污染核的遍历系数与从污染数据中学习的难度之间有何关系?
主要发现
- 本文建立了基于污染核遍历系数的、关于从污染数据中学习风险的下界。
- 通过广义无偏估计方法,提供了风险的上界,该方法扩展了[30]和[24]中的先前工作。
- 对于对称标签噪声,(0-1损失, 污染核)对在η = max((1+σ₋₁−σ₁)/(1−σ₋₁−σ₁))²下具有η-相容性,从而在Bernstein条件下可实现快速率。
- 若干净数据满足常数为K的Bernstein条件,且污染为η-相容,则污染数据也满足常数为ηK的Bernstein条件。
- 该框架可通过界定相对学习难度,实现对不同数据类型的比较——例如,n个干净样本与n₁个噪声标签和n₂个部分标签的比较。
- 结果表明,在某些条件下(如可分离数据与0-1损失),即使存在标签污染,快速率仍能保持,为高效利用噪声数据提供了理论依据。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。