Skip to main content
QUICK REVIEW

[论文解读] Compression, Generalization and Learning

Marco C. Campi, Simone Garatti|arXiv (Cornell University)|Jan 30, 2023
Machine Learning and Algorithms被引用 5
一句话总结

该论文建立了一个新颖的理论框架,将压缩集大小与压缩过程中变化的概率联系起来,无需分布假设即可提供有限样本的误分类风险界。研究表明,压缩集的基数始终一致地估计变化概率,通过偏好条件可实现紧密且无偏的界。

ABSTRACT

A compression function is a map that slims down an observational set into a subset of reduced size, while preserving its informational content. In multiple applications, the condition that one new observation makes the compressed set change is interpreted that this observation brings in extra information and, in learning theory, this corresponds to misclassification, or misprediction. In this paper, we lay the foundations of a new theory that allows one to keep control on the probability of change of compression (which maps into the statistical "risk" in learning applications). Under suitable conditions, the cardinality of the compressed set is shown to be a consistent estimator of the probability of change of compression (without any upper limit on the size of the compressed set); moreover, unprecedentedly tight finite-sample bounds to evaluate the probability of change of compression are obtained under a generally applicable condition of preference. All results are usable in a fully agnostic setup, i.e., without requiring any a priori knowledge on the probability distribution of the observations. Not only these results offer a valid support to develop trust in observation-driven methodologies, they also play a fundamental role in learning techniques as a tool for hyper-parameter tuning.

研究动机与目标

  • 开发一种理论,以在不了解数据分布的情况下控制压缩函数中变化的概率。
  • 确立压缩集的基数作为变化概率的一致估计量,即使在集合大小无界的情况下亦成立。
  • 通过一般偏好条件推导前所未有的紧致有限样本界,用于误分类风险。
  • 使压缩理论在监督学习和无监督学习以及其他领域中能够以完全无偏的方式应用。
  • 通过严格的统计控制,支持对基于观测的方法和超参数调优的信任。

提出的方法

  • 引入一种压缩函数,将观测集合映射为保留信息内容的子集。
  • 将压缩的变化概率定义为新观测改变压缩集的可能性。
  • 利用偏好条件推导出与数据分布无关的有限样本变化概率界。
  • 应用涉及多重集和概率测度的测度论构造,以建模压缩集的演化过程。
  • 通过测度逼近序列的极限论证证明紧致界,利用非递减函数和质量重分配。
  • 确立压缩集的基数始终一致地估计变化风险,即使在无大小上界的情况下亦成立。

实验结果

研究问题

  • RQ1在不假设数据分布的前提下,如何在有限样本中界定压缩函数中变化的概率?
  • RQ2压缩集的大小能否作为压缩中变化概率的一致估计量?
  • RQ3在何种条件下可通过压缩推导出紧致且无偏的有限样本误分类风险界?
  • RQ4偏好条件如何使界比统计学习理论中现有方法更紧密?
  • RQ5所提出的框架在何种方式下可支持无分布假设下的学习系统中的信任与超参数调优?

主要发现

  • 即使集合大小无界,压缩集的基数仍是压缩变化概率的一致估计量。
  • 在一般偏好条件下推导出有限样本变化概率的界,其保证比以往方法更紧密。
  • 该界具有无偏性,无需对底层数据分布的先验知识,因此可广泛应用于各类学习场景。
  • 该框架可通过仅基于压缩的分析,在监督学习和无监督学习情境中实现风险评估。
  • 理论结果支持将压缩集大小作为统计风险的代理变量,从而促进学习系统中的超参数调优。
  • 证明技术通过构造保持约束并收敛至最优界的逼近测度序列,确保了界的一致紧致性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。