Skip to main content
QUICK REVIEW

[论文解读] Inference under Information Constraints I: Lower Bounds from Chi-Square Contraction

Jayadev Acharya, Clément L. Canonne|arXiv (Cornell University)|Dec 30, 2018
Privacy-Preserving Technologies in Data被引用 5
一句话总结

本文通过一种新颖的卡方收缩框架,建立了在信息约束下学习和测试离散分布的紧致样本复杂度下界。它引入了局部和解耦卡方波动,以表征通信限制或局部微差隐私等信息约束如何降低统计可区分性,揭示了在相同约束下,私有随机协议的样本复杂度高于公共随机协议。

ABSTRACT

Multiple players are each given one independent sample, about which they can only provide limited information to a central referee. Each player is allowed to describe its observed sample to the referee using a channel from a family of channels $\\mathcal{W}$, which can be instantiated to capture both the communication- and privacy-constrained settings and beyond. The referee uses the messages from players to solve an inference problem for the unknown distribution that generated the samples. We derive lower bounds for sample complexity of learning and testing discrete distributions in this information-constrained setting. Underlying our bounds is a characterization of the contraction in chi-square distances between the observed distributions of the samples when information constraints are placed. This contraction is captured in a local neighborhood in terms of chi-square and decoupled chi-square fluctuations of a given channel, two quantities we introduce. The former captures the average distance between distributions of channel output for two product distributions on the input, and the latter for a product distribution and a mixture of product distribution on the input. Our bounds are tight for both public- and private-coin protocols. Interestingly, the sample complexity of testing is order-wise higher when restricted to private-coin protocols.

研究动机与目标

  • 建立在局部信息约束(如通信限制和局部微差隐私)下统计推断的样本复杂度基本下界。
  • 通过将参与者的消息建模为固定族中的随机信道,构建一个分析信息约束推断的一般性框架。
  • 利用卡方距离收缩表征信息约束导致的统计可区分性降低。
  • 为公共随机和私有随机协议下的学习与测试任务推导出紧致边界。
  • 证明在相同约束下,私有随机协议的样本复杂度比公共随机协议高阶。

提出的方法

  • 引入局部卡方波动的概念,用于度量将信道应用于两个乘积输入分布时,输出分布之间平均卡方距离。
  • 定义解耦卡方波动,以量化在信道作用下,乘积分布与乘积分布混合之间的距离。
  • 利用这些波动来界定施加信息约束后输出分布的卡方收缩。
  • 应用卡方散度的数据处理不等式,将输入扰动与在信道约束下输出的可区分性关联起来。
  • 通过将最小可实现卡方距离与可靠推断所需的样本量关联,推导出样本复杂度的下界。
  • 通过将前者建模为共享随机性、后者建模为信道的凸组合,区分公共随机与私有随机协议。

实验结果

研究问题

  • RQ1信息约束(如通信受限或局部微差隐私)如何影响学习离散分布的样本复杂度?
  • RQ2在分布式推断中,信息约束与统计可区分性之间的基本权衡是什么?
  • RQ3在信息约束下,私有随机与公共随机协议在分布测试中的样本复杂度有何比较?
  • RQ4卡方收缩能否用于推导信息约束环境下学习与测试的紧致下界?
  • RQ5局部与解耦卡方波动在表征信道约束导致的信息损失方面起什么作用?

主要发现

  • 本文通过卡方收缩方法,为在信息约束下学习离散分布的样本复杂度建立了紧致下界,表明该下界依赖于信道族的局部卡方波动。
  • 在分布测试中,即使信道族相同,私有随机协议的样本复杂度也比公共随机协议高阶。
  • 卡方收缩边界通过一种新颖的表征方法推导得出,其中涉及局部与解耦卡方波动,用于捕捉信道作用下平均与混合输入输出的偏离程度。
  • 通过构造分布族的近似 ε-扰动并证明:只有当样本量足够大时,输出分布才能保持统计可区分,从而推导出测试的下界。
  • 分析证明,在私有随机设置下,对于信道族中的任意有效信道选择,输出分布之间的最小卡方距离被下界为常数 c = log(14401/14400)。
  • 该框架在学习与测试任务中均达到紧致性,其边界与通信和隐私约束环境下文献中已知的上界一致。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。