[论文解读] Provably Strict Generalisation Benefit for Equivariant Models
本文通过严格量化在紧致群作用下模型被约束为等变时测试误差的改进,首次为等变线性模型提供了可证明的非零泛化优势。关键结果表明,泛化差距取决于对称性显著性(通过特征标内积衡量),并在插值阈值处发散,展示了在特定分布假设下,与标准模型相比存在严格且可量化的优越性。
It is widely believed that engineering a model to be invariant/equivariant improves generalisation. Despite the growing popularity of this approach, a precise characterisation of the generalisation benefit is lacking. By considering the simplest case of linear models, this paper provides the first provably non-zero improvement in generalisation for invariant/equivariant models when the target distribution is invariant/equivariant with respect to a compact group. Moreover, our work reveals an interesting relationship between generalisation, the number of training examples and properties of the group action. Our results rest on an observation of the structure of function spaces under averaging operators which, along with its consequences for feature averaging, may be of independent interest.
研究动机与目标
- 正式建立线性设定下等变模型的可证明、非零泛化优势。
- 描述群对称性、训练数据规模与表示理论如何共同影响泛化性能。
- 解决关于等变性是否在最坏情况界限之外真正改善泛化性能的理论模糊性。
- 基于平均算子和表示理论,构建严谨框架以分析函数空间结构。
- 基于理论原则,推导出用于训练不变/等变神经网络的实用洞见。
提出的方法
- 使用最小 L2-范数最小二乘解来建模过参数化线性模型中梯度下降的隐式偏差。
- 应用正交群表示和特征标理论,计算表示的内积,以捕捉对称性显著性。
- 采用平均算子将函数空间分解为不变、等变和反对称子空间。
- 利用随机矩阵理论和各向同性设计假设,推导出期望泛化差距的精确表达式。
- 引入矩阵值函数 $ J_{/mathcal{G}} $ 以编码群作用特性及其对泛化的影响。
- 使用投影算子 $ \Psi_{\mathcal{G}}^{\perp} $ 隔离违反等变性的模型分量,量化其对测试误差的贡献。
实验结果
研究问题
- RQ1在线性模型中强制实施等变性是否能导致与标准模型相比,泛化性能得到可证明的非零改进?
- RQ2训练样本数量 $ n $ 与群对称性如何相互作用以影响泛化性能?
- RQ3特征标内积 $ (\chi_\psi|\chi_\phi) $ 在决定泛化优势大小方面起什么作用?
- RQ4为何泛化差距在插值阈值 $ n \in [d-1, d+1] $ 处发散,这与双下降现象如何一致?
- RQ5线性模型中的洞见能否用于设计更好的不变/等变神经网络训练方法?
主要发现
- 期望泛化差距为 $ \mathbb{E}[\Delta(f,f^\prime)] = \text{Var}[\xi] \cdot r(n,d) \cdot (dk - (\chi_\psi|\chi_\phi)) + \mathcal{E}_{\mathcal{G}}(n,d) $,其中 $ r(n,d) $ 捕获数据制度依赖性。
- 当 $ dk - (\chi_\psi|\chi_\phi) > 0 $ 时,泛化优势严格为正,即当任务具有非平凡对称结构时。
- $ r(n,d) $ 在 $ n \in [d-1, d+1] $ 处发散,与双下降现象一致,表明在插值附近泛化增益增加。
- 项 $ dk - (\chi_\psi|\chi_\phi) $ 表示在群平均下消失的线性映射空间的维数,直接将对称性与泛化增益联系起来。
- 当 $ n \geq d $ 时,无噪声泛化差距 $ \mathcal{E}_{\mathcal{G}}(n,d) $ 消失,表明过参数化模型可恢复真实等变函数。
- 推导出的 $ \mathbb{E}[\|\Psi_{\mathcal{G}}^{\perp}(X^\top X)^+ X^\top X \Theta\|_F^2] $ 表达式,精确建立了模型偏离等变性与测试误差之间的联系。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。