Skip to main content
QUICK REVIEW

[论文解读] On the exact relationship between the denoising function and the data distribution

Heikki Arponen, Matti Herranen|arXiv (Cornell University)|Sep 6, 2017
NMR spectroscopy and applications参考文献 3被引用 3
一句话总结

该论文在加性高斯噪声下建立了最优去噪函数与数据分布之间的精确、可逆关系,表明去噪函数在数学上等价于受损数据分布的得分函数。关键贡献在于将先前仅在小噪声极限下成立的结果推广至任意噪声水平,证明了通过对数概率密度的梯度,去噪过程隐式地建模了完整的数据流形结构。

ABSTRACT

We prove an exact relationship between the optimal denoising function and the data distribution in the case of additive Gaussian noise, showing that denoising implicitly models the structure of data allowing it to be exploited in the unsupervised learning of representations. This result generalizes a known relationship [2], which is valid only in the limit of small corruption noise.

研究动机与目标

  • 建立在加性高斯噪声下最优去噪函数与底层数据分布之间精确的数学关系。
  • 将先前仅在小干扰噪声极限下有效的理论结果推广至一般情况。
  • 证明通过建模受损数据分布的得分函数,去噪学习隐式捕捉了数据流形的完整结构。
  • 表明去噪函数可逆,从而恢复数据分布,实现精确表征学习。
  • 为使用去噪自编码器作为无监督表征学习的合理方法奠定理论基础。

提出的方法

  • 将最优去噪函数推导为最小均方误差(MMSE)估计器,其表达式为 $ g^*(\tilde{x}) = \mathbb{E}[x|\tilde{x}] $。
  • 利用贝叶斯法则,将条件期望表示为先验数据分布 $ p(x) $ 和损坏似然 $ p(\tilde{x}|x) $ 的函数。
  • 应用加性高斯噪声的特定形式 $ \tilde{x} = x + \sigma_n \epsilon $,其中 $ \epsilon \sim \mathcal{N}(0,I) $,以推导出精确关系。
  • 推导出关键恒等式 $ x p(\tilde{x}|x) = \tilde{x} p(\tilde{x}|x) + \sigma_n^2 \nabla_{\tilde{x}} p(\tilde{x}|x) $,该式将干净输入与受损观测联系起来。
  • 结合对 $ x $ 的边缘化,证明 $ g^*(\tilde{x}) = \tilde{x} + \sigma_n^2 \nabla_{\tilde{x}} \log p(\tilde{x}) $,其中 $ p(\tilde{x}) $ 为受损数据分布。
  • 通过围道积分证明该关系可逆,从而可从 $ g^* $ 重建 $ p(\tilde{x}) $,并进一步通过去卷积恢复 $ p(x) $。

实验结果

研究问题

  • RQ1在加性高斯噪声下,最优去噪函数与数据分布之间的确切数学关系是什么?
  • RQ2该关系如何推广先前仅在小噪声极限下成立的结果?
  • RQ3能否利用去噪函数恢复完整的数据分布,包括其流形结构?
  • RQ4去噪函数与数据分布之间的映射是否可逆?
  • RQ5在深度学习中,使用去噪作为表征学习方法的理论依据是什么?

主要发现

  • 最优去噪函数精确为 $ g^*(\tilde{x}) = \tilde{x} + \sigma_n^2 \nabla_{\tilde{x}} \log p(\tilde{x}) $,其中 $ p(\tilde{x}) $ 为受损数据分布。
  • 该结果将已知的小噪声近似 $ g^*(\tilde{x}) = \tilde{x} + \sigma_n^2 \nabla_{\tilde{x}} \log p_X(\tilde{x}) + o(\sigma_n^2) $ 推广至任意噪声水平。
  • 去噪函数与数据分布之间的关系是可逆的:$ p(\tilde{x}) $ 可通过围道积分从 $ g^* $ 恢复。
  • 去噪函数通过受损分布的得分函数编码了数据流形的完整结构。
  • 映射 $ p(x) \longleftrightarrow p(\tilde{x}) \longleftrightarrow g^*(\tilde{x}) $ 是精确且可逆的,证明了去噪学习了完整的数据分布。
  • 该结果表明,去噪自编码器能够学习超越简单条件期望的复杂数据结构,与标准回归方法不同。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。