[论文解读] Locally Private Hypothesis Testing
本文在本地差分隐私模型中引入了差分私有假设检验,分析了对称(随机响应)与非对称机制。研究建立了在本地隐私下的可行最大似然估计,利用非对称机制推导出更紧致的独立性与拟合优度检验的样本复杂度边界,并提出一种基于卡方检验的测试方法,经实证验证表明其效率优于对称机制。
We initiate the study of differentially private hypothesis testing in the local-model, under both the standard (symmetric) randomized-response mechanism (Warner, 1965, Kasiviswanathan et al, 2008) and the newer (non-symmetric) mechanisms (Bassily and Smith, 2015, Bassily et al, 2017). First, we study the general framework of mapping each user's type into a signal and show that the problem of finding the maximum-likelihood distribution over the signals is feasible. Then we discuss the randomized-response mechanism and show that, in essence, it maps the null- and alternative-hypotheses onto new sets, an affine translation of the original sets. We then give sample complexity bounds for identity and independence testing under randomized-response. We then move to the newer non-symmetric mechanisms and show that there too the problem of finding the maximum-likelihood distribution is feasible. Under the mechanism of Bassily et al (2007) we give identity and independence testers with better sample complexity than the testers in the symmetric case, and we also propose a $χ^2$-based identity tester which we investigate empirically.
研究动机与目标
- 研究用户独立应用私有机制的本地模型中的差分私有假设检验。
- 分析在本地隐私机制下最大似然估计的可行性,包括对称与非对称设计。
- 推导在本地差分隐私下的拟合优度检验与独立性检验的样本复杂度边界。
- 通过利用随类型空间大小对数级增长的非对称机制,改进现有对称机制。
- 研究基于非对称机制的卡方检验方法在拟合优度检验中的实证性能。
提出的方法
- 将本地差分隐私建模为一种信号方案,通过随机机制将用户类型映射到信号,确保 ε-差分隐私。
- 分析对称机制(如随机响应)作为索引无关的映射,表明其在原假设与备择假设集合上诱导仿射平移。
- 引入非对称机制(如 Bassily 等 [BNST17]),利用公开的、用户特定的映射将类型映射到信号,实现类型空间大小的对数级缩放。
- 推导在对称与非对称机制下对信号分布的最大似然估计程序。
- 提出一种基于卡方检验的统计量用于在非对称机制下的拟合优度检验,并实证评估其在原假设与备择假设下的分布。
- 采用二分查找与经验拒绝概率估计样本复杂度,拟合出推测形式 $ T^{1.5}/(\alpha^2 \epsilon^2) $。
实验结果
研究问题
- RQ1在本地差分隐私机制(包括非对称机制)下,最大似然估计是否能够高效计算?
- RQ2在本地隐私下,拟合优度与独立性检验的样本复杂度在对称与非对称机制之间有何差异?
- RQ3基于卡方检验的统计量是否可在本地模型中有效用于拟合优度检验,且能否有效区分原假设与备择假设?
- RQ4本地私有测试器在高概率下拒绝原假设所需的实证样本复杂度是多少?
- RQ5所提出的卡方检验统计量在原假设下是否收敛到卡方分布,或在实证中可被明确区分?
主要发现
- 在对称与非对称本地隐私机制下,对信号分布的最大似然分布均可高效计算。
- 非对称机制(如 [BNST17])在拟合优度与独立性检验中,其测试器的样本复杂度优于对称机制。
- 实证结果表明样本复杂度约为 $ T^{1.5}/(\alpha^2 \epsilon^2) $,其中指数 $ c_T \approx 1.49 $,$ c_\alpha \approx -1.93 $,$ c_\epsilon \approx -1.90 $。
- 所提出的基于卡方检验的拟合优度测试器在原假设与备择假设下的分布存在显著差异,表明其具有潜在效用,尽管其分布不完全匹配理论卡方分布。
- 在非对称机制下,边缘分布的估计量并非相互独立,导致无法直接对卡方统计量求和,此问题仍为开放问题。
- 实证评估证实,测试统计量 $ Q(\boldsymbol{\theta}) = n\sum_x \frac{(\frac{1}{2\eta}\theta(x) - \bar{\theta}(x))^2}{\bar{\theta}(x)} $ 在不同假设下具有可区分性,支持其作为候选测试统计量的使用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。