[Paper Review] Protection Against Reconstruction and Its Applications in Private Federated Learning
The paper proposes a privacy framework focusing on protection against reconstruction for locally private federated learning, develops minimax-optimal local privacy mechanisms for high-dimensional data, and demonstrates practical large-scale private model training with limited utility loss.
In large-scale statistical learning, data collection and model fitting are moving increasingly toward peripheral devices---phones, watches, fitness trackers---away from centralized data collection. Concomitant with this rise in decentralized data are increasing challenges of maintaining privacy while allowing enough information to fit accurate, useful statistical models. This motivates local notions of privacy---most significantly, local differential privacy, which provides strong protections against sensitive data disclosures---where data is obfuscated before a statistician or learner can even observe it, providing strong protections to individuals' data. Yet local privacy as traditionally employed may prove too stringent for practical use, especially in modern high-dimensional statistical and machine learning problems. Consequently, we revisit the types of disclosures and adversaries against which we provide protections, considering adversaries with limited prior information and ensuring that with high probability, ensuring they cannot reconstruct an individual's data within useful tolerances. By reconceptualizing these protections, we allow more useful data release---large privacy parameters in local differential privacy---and we design new (minimax) optimal locally differentially private mechanisms for statistical learning problems for \emph{all} privacy levels. We thus present practicable approaches to large-scale locally private model training that were previously impossible, showing theoretically and empirically that we can fit large-scale image classification and language models with little degradation in utility.
Motivation & Objective
- Motivate local privacy in decentralized data settings and address reconstruction risks in federated learning.
- Propose a refined threat model focusing on reconstruction by curious onlookers with limited prior information.
- Develop minimax-optimal locally private mechanisms for high-dimensional vectors across all privacy levels (ε ≤ d).
- Demonstrate practical, large-scale private model training with acceptable utility degradation.
- Provide empirical results in image classification and language modeling under local privacy protections.
Proposed method
- Define a reconstruction-protection privacy model where an adversary with limited prior information attempts to reconstruct data from privatized outputs.
- Introduce ε-local differential privacy and Reconstruction Breach concepts to quantify protection against data reconstruction.
- Develop new minimax-optimal privatization mechanisms for high-dimensional vectors in unit balls.
- Analyze asymptotic behavior of a stochastic-gradient-based private learning scheme under these mechanisms.
- Embed the local privacy layer within a broader central differential privacy framework for end-to-end protection.
- Demonstrate a prototypical private federated learning system with private updates and centralized privacy accounting.
Experimental results
Research questions
- RQ1How can we protect against accurate reconstruction of private data in locally private federated learning settings?
- RQ2What are minimax-optimal locally private mechanisms for high-dimensional data across ε in [0, d]?
- RQ3Can large-scale image and language models be trained privately with limited degradation in utility?
- RQ4How does reconstructions protection interact with centralized differential privacy in a federated setup?
Key findings
- Locally private mechanisms with ε up to d achieve minimax-optimal performance across privacy levels.
- Reconstruction breaches can be tightly bounded under diffuse priors, with protections improving as ε increases and priors become more informative.
- The proposed framework yields practical procedures enabling large-scale locally private model training with little utility loss compared to non-private baselines.
- Experiments indicate feasibility of private federated learning for image classification and language models under the proposed protections.
- The combination of local reconstruction protection with centralized DP maintains strong privacy while supporting scalable distributed learning.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.