Skip to main content
QUICK REVIEW

[论文解读] Universality of Computational Lower Bounds for Submatrix Detection

Matthew Brennan, Guy Bresler|arXiv (Cornell University)|Feb 19, 2019
Statistical Methods and Inference参考文献 40被引用 10
一句话总结

该论文通过从植 clique 猜想出发的平均案例归约,建立了子矩阵检测中计算下界的普遍性原理,表明统计-计算间隙普遍依赖于分布 P 与 Q 之间的 KL 散度。关键结果是将计算障碍紧致表征为子矩阵大小 k 与 KL 散度之间的权衡,推广了针对高斯双聚类和植稠密子图的先前结果。

ABSTRACT

In the general submatrix detection problem, the task is to detect the presence of a small $k imes k$ submatrix with entries sampled from a distribution $\mathcal{P}$ in an $n imes n$ matrix of samples from $\mathcal{Q}$. This formulation includes a number of well-studied problems, such as biclustering when $\mathcal{P}$ and $\mathcal{Q}$ are Gaussians and the planted dense subgraph formulation of community detection when the submatrix is a principal minor and $\mathcal{P}$ and $\mathcal{Q}$ are Bernoulli random variables. These problems all seem to exhibit a universal phenomenon: there is a statistical-computational gap depending on $\mathcal{P}$ and $\mathcal{Q}$ between the minimum $k$ at which this task can be solved and the minimum $k$ at which it can be solved in polynomial time. Our main result is to tightly characterize this computational barrier as a tradeoff between $k$ and the KL divergences between $\mathcal{P}$ and $\mathcal{Q}$ through average-case reductions from the planted clique conjecture. These computational lower bounds hold given mild assumptions on $\mathcal{P}$ and $\mathcal{Q}$ arising naturally from classical binary hypothesis testing. Our results recover and generalize the planted clique lower bounds for Gaussian biclustering in Ma-Wu (2015) and Brennan et al. (2018) and for the sparse and general regimes of planted dense subgraph in Hajek et al. (2015) and Brennan et al. (2018). This yields the first universality principle for computational lower bounds obtained through average-case reductions.

研究动机与目标

  • 确定子矩阵检测中的统计-计算间隙是否在不同分布 P 和 Q 下具有普遍性。
  • 以 P 与 Q 之间的 KL 散度而非特定问题参数来表征计算障碍。
  • 开发适用于广泛子矩阵检测问题的平均案例归约通用技术。
  • 恢复并推广基于植 clique 的高斯双聚类和植稠密子图的现有下界。
  • 通过平均案例归约建立计算下界普遍性原理,这是该领域的一项新贡献。

提出的方法

  • 引入多变量拒绝核(MRK),实现算法测度变换,并在保持最优 KL 散度权衡的同时提升子矩阵。
  • 提出一种将图邻接矩阵嵌入更大矩阵的主子式中的技术,解决了分布嵌入中的对角线和支撑匹配问题。
  • 在总变差距离下使用平均案例归约,将植 clique 问题与在二元假设检验的温和假设下的一般子矩阵检测联系起来。
  • 将植 clique 猜想作为难解性假设,推导出子矩阵检测的紧致计算下界。
  • 采用基于对数似然比的分析,推导出 MRK 实现 KL 散度最优权衡的条件。
  • 基于分布特性与 KL 散度缩放关系,定义了三个普遍性类 UC-A、UC-B 和 UC-C。

实验结果

研究问题

  • RQ1子矩阵检测中的统计-计算间隙是否在不同分布对 (P, Q) 间具有普遍性?
  • RQ2子矩阵检测的计算下界能否由 P 与 Q 之间的 KL 散度紧致表征,而独立于具体问题结构?
  • RQ3哪些通用技术可实现从植 clique 到任意 (P, Q) 对的子矩阵检测的平均案例归约?
  • RQ4是否存在自然的 (P, Q) 对普遍性类,其相图特征超出本文识别的范围?
  • RQ5检测的计算障碍与子矩阵恢复的计算障碍如何比较,能否通过归约建立联系?

主要发现

  • 子矩阵检测的计算障碍被普遍表征为子矩阵大小 k 与 P 和 Q 之间 KL 散度之间的权衡,通过平均案例归约推导出紧致边界。
  • 所提出的多变量拒绝核(MRK)技术实现了最优 KL 散度权衡,其在实现紧致下界方面优于逐元素归约方法。
  • 主子式嵌入技术成功解决了图到矩阵归约中对角线元素和行/列支撑匹配带来的分布不一致问题。
  • 结果恢复并推广了针对高斯双聚类(MW 15;BBH 18)和植稠密子图在稀疏与一般情形下的植 clique 下界(HWX 15;BBH 18)。
  • 识别出三个普遍性类(UC-A、UC-B、UC-C),所有测试的分布对(如高斯、伯努利)均落入 UC-C 类,其中 KL 散度按 Θ(n^{-α}) 或 Θ(n^{-2γ+α}) 缩放。
  • 子矩阵检测的信息论极限通过信息理论下界表征,而在植 clique 猜想下,计算极限被证明严格高于该极限。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。