[论文解读] Protection Against Reconstruction and Its Applications in Private Federated Learning
论文提出一种专注于对本地隐私联邦学习的重建保护的隐私框架,开发针对高维数据的极小极大值最优本地隐私机制,并证明在有限效用损失的情况下进行实际的大规模私有模型训练。
In large-scale statistical learning, data collection and model fitting are moving increasingly toward peripheral devices---phones, watches, fitness trackers---away from centralized data collection. Concomitant with this rise in decentralized data are increasing challenges of maintaining privacy while allowing enough information to fit accurate, useful statistical models. This motivates local notions of privacy---most significantly, local differential privacy, which provides strong protections against sensitive data disclosures---where data is obfuscated before a statistician or learner can even observe it, providing strong protections to individuals' data. Yet local privacy as traditionally employed may prove too stringent for practical use, especially in modern high-dimensional statistical and machine learning problems. Consequently, we revisit the types of disclosures and adversaries against which we provide protections, considering adversaries with limited prior information and ensuring that with high probability, ensuring they cannot reconstruct an individual's data within useful tolerances. By reconceptualizing these protections, we allow more useful data release---large privacy parameters in local differential privacy---and we design new (minimax) optimal locally differentially private mechanisms for statistical learning problems for \emph{all} privacy levels. We thus present practicable approaches to large-scale locally private model training that were previously impossible, showing theoretically and empirically that we can fit large-scale image classification and language models with little degradation in utility.
研究动机与目标
- 在去中心化数据环境中激发本地隐私的动机并解决联邦学习中的重建风险。
- 提出一个聚焦于好奇旁观者在有限先验信息下的重建的 refined threat model。
- 开发在所有隐私级别(ε ≤ d)下针对高维向量的极小极大值最优本地化隐私机制。
- 证明在可接受的效用下降下进行实际的、大规模私有模型训练。
- 在本地隐私保护下提供在图像分类和语言建模方面的实证结果。
提出的方法
- 定义一个重建保护隐私模型,其中具有限先验信息的对手试图从私有化输出中重建数据。
- 引入 ε-local differential privacy and Reconstruction Breach 的概念以量化对数据重建的保护。
- 开发针对单位球面高维向量的新型极小极大值最优 privatization 机制。
- 分析在这些机制下基于随机梯度的私有学习方案的渐近行为。
- 在端到端保护的更广泛中心化差分隐私框架中嵌入本地隐私层。
- 演示一个原型私有联邦学习系统,具有私有更新和集中隐私审计。
实验结果
研究问题
- RQ1我们如何在本地隐私联邦学习设置中防止对私有数据的精准重建?
- RQ2在 ε ∈ [0, d] 范围内高维数据的极小极大值最优本地隐私机制是什么?
- RQ3在有限的效用下降下,是否能够大规模地私有训练图像和语言模型?
- RQ4在联邦架构中,重建保护如何与集中式差分隐私相互作用?
主要发现
- 在 ε 高达 d 的本地隐私机制在各隐私水平上实现了极小极大值最优的性能。
- 在分散先验下,重建违反可以被严格限定,且随着 ε 增加和先验信息更充分,保护效果提升。
- 所提出的框架能够提供实用的程序,使得在本地隐私保护下进行大规模模型训练且与非私有基线相比有很小的效用损失。
- 实验表明在所提出的保护下,私有联邦学习在图像分类和语言模型方面具有可行性。
- 将本地重建保护与中心化 DP 的结合在支持可扩展分布式学习的同时保持强隐私性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。