Skip to main content
QUICK REVIEW

[论文解读] Exponential Series Approaches for Nonparametric Graphical Models

Eric Janofsky|arXiv (Cornell University)|Jun 11, 2015
Statistical Methods and Inference参考文献 69被引用 6
一句话总结

本文提出了一种基于正则化最大似然估计和得分匹配的指数级数方法,用于非参数图模型,实现了高维密度估计中的稀疏性和风险一致性。在弱光滑性假设下,该方法建立了参数和模型选择一致性的理论保证,并在合成数据和真实世界MEG数据上进行了经验验证。

ABSTRACT

This thesis studies high-dimensional, continuous-valued pairwise Markov Random Fields. We are particularly interested in approximating pairwise densities whose logarithm belongs to a Sobolev space. For this problem we propose the method of exponential series [Crain, 1974; Barron and Sheu, 1991], which approximates the log density by a finite- dimensional exponential family with the number of sufficient statistics increasing with the sample size. We consider two approaches to estimating these models. The first is regularized maximum likelihood. This involves optimizing the sum of the log-likelihood of the data and a sparsity-inducing regularizer. We provide consistency and edge selection guarantees for this method. We then propose a variational approximation to the likelihood based on tree- reweighted, nonparametric message passing. We then consider estimation using regularized score matching. This approach uses an alternative scoring rule to the log-likelihood, which obviates the need to compute the normalizing constant of the distribution. For general continuous-valued exponential families, we provide parameter and edge consistency results. We then describe results for model selection in the nonparametric pairwise model using exponential series. The regularized score matching problem is shown to be a convex program; we provide scalable algorithms based on consensus Alternating Direction Method of Multipliers (ADMM, [Boyd et al., 2011]) and Coordinate-wise Descent. We compare our method to others in the literature as well as the aforementioned TRW estimator using simulated data.

研究动机与目标

  • 通过在高维数据中建模成对依赖关系,解决非参数密度估计中的维度灾难问题。
  • 开发无需对联合密度做参数假设的理论基础坚实的无向图模型估计方法。
  • 通过正则化和函数逼近,在高维设置下确保稀疏性和模型选择一致性。
  • 在真实密度的弱光滑性条件下,提供风险一致性和参数估计的理论保证。
  • 通过新型算法(如函数消息传递和一致性ADMM)实现实际可计算的估计。

提出的方法

  • 在希尔伯特空间中使用指数级数展开来近似非参数成对密度,利用正交基函数(如勒让德多项式)。
  • 应用带ℓ1型惩罚的正则化最大似然估计,以在估计的图模型中诱导稀疏性。
  • 引入树重加权变分似然近似(TRW),以提高高维下似然函数的可处理性。
  • 采用函数消息传递和一致性ADMM算法,高效优化正则化似然和得分匹配目标。
  • 通过原始-对偶见证条件和浓度不等式推导理论界,建立参数和模型选择一致性的理论基础。
  • 使用得分匹配并采用Hyvärinen型评分规则,避免归一化常数,从而在非指数族设定下实现估计。

实验结果

研究问题

  • RQ1在弱光滑性假设下,指数级数近似能否在非参数图模型中实现风险一致性和模型选择一致性?
  • RQ2对函数系数施加ℓ1惩罚的正则化如何影响高维成对密度估计中的稀疏性和估计精度?
  • RQ3在非参数设定下,一致恢复参数和图结构所需的理论样本复杂度是多少?
  • RQ4在计算效率和边选择准确性方面,函数消息传递和一致性ADMM相较于标准方法(如glasso)表现如何?
  • RQ5使用非参数基展开的得分匹配能否在估计复杂非高斯依赖结构方面优于正则化MLE?

主要发现

  • 在弱光滑性条件下,所提出的带指数级数近似的正则化MLE可实现 $ O\big(\frac{\text{polylog}(d)}{n}\big) $ 的风险一致性收敛速率。
  • 当调参 $ \rho_n $ 满足 $ \frac{\rho_n}{\rho^*} \to 0 $ 时,可建立模型选择一致性,确保正确包含或排除边。
  • 对于非参数成对模型,当真实密度属于阶数为 $ r $ 的Sobolev空间时,该方法可实现 $ \tilde{O}(n^{-\frac{2r-1}{2r+13}}) $ 的估计误差率。
  • 实证结果表明,基于TRW的SKEPTIC估计器在合成高斯数据和拷贝拉数据上的边选择性能优于glasso和TRW,尤其在高维设置下表现更优。
  • 在MEG数据上,所提方法恢复了具有生物学合理性的脑连接模式,其边选择性能优于glasso和TRW。
  • QUASR算法(基于一致性ADMM)的正则化路径在不同样本大小下均表现出稳定收敛和准确的图恢复能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。