[Paper Review] Privacy-Preserving Machine Learning: Methods, Challenges and Directions
A systematic review of privacy-preserving ML (PPML) that introduces the PGU triad (Phase, Guarantee, Utility) to evaluate PPML solutions, and outlines a taxonomy, challenges, and future directions.
Machine learning (ML) is increasingly being adopted in a wide variety of application domains. Usually, a well-performing ML model relies on a large volume of training data and high-powered computational resources. Such a need for and the use of huge volumes of data raise serious privacy concerns because of the potential risks of leakage of highly privacy-sensitive information; further, the evolving regulatory environments that increasingly restrict access to and use of privacy-sensitive data add significant challenges to fully benefiting from the power of ML for data-driven applications. A trained ML model may also be vulnerable to adversarial attacks such as membership, attribute, or property inference attacks and model inversion attacks. Hence, well-designed privacy-preserving ML (PPML) solutions are critically needed for many emerging applications. Increasingly, significant research efforts from both academia and industry can be seen in PPML areas that aim toward integrating privacy-preserving techniques into ML pipeline or specific algorithms, or designing various PPML architectures. In particular, existing PPML research cross-cut ML, systems and applications design, as well as security and privacy areas; hence, there is a critical need to understand state-of-the-art research, related challenges and a research roadmap for future research in PPML area. In this paper, we systematically review and summarize existing privacy-preserving approaches and propose a Phase, Guarantee, and Utility (PGU) triad based model to understand and guide the evaluation of various PPML solutions by decomposing their privacy-preserving functionalities. We discuss the unique characteristics and challenges of PPML and outline possible research directions that leverage as well as benefit multiple research communities such as ML, distributed systems, security and privacy.
Motivation & Objective
- Motivate the need for PPML due to privacy risks and regulatory constraints in ML pipelines.
- Propose a holistic framework (PGU) to evaluate PPML approaches across phases, guarantees, and utility.
- Classify PPML solutions into data publishing, data processing, architectural, and hybrid categories.
- Analyze privacy guarantees from object-oriented and pipeline-oriented perspectives.
Proposed method
- Propose the PGU (Phase, Guarantee, Utility) triad to decompose PPML functionalities.
- Map PPML solutions to privacy-preserving phases: data preparation, model generation, serving, and inference.
- Differentiate object-oriented privacy guarantees (input data and model weights) from pipeline-oriented guarantees (local, global, full-chain privacy).
- Classify PPML techniques into data publishing, data processing, architectural, and hybrid approaches, and assess their utility impacts.
- Discuss challenges and directions for measuring privacy, attack/defense strategies, and efficiency considerations.
Experimental results
Research questions
- RQ1What are the core privacy-preserving functionalities that PPML approaches provide across ML pipelines?
- RQ2How can the PGU framework be used to evaluate the strength and scope of privacy guarantees in PPML solutions?
- RQ3What taxonomy best captures the technical approaches and their impact on utility in PPML?
- RQ4What are the open challenges and promising directions for future PPML research?
Key findings
- PPML solutions are diverse and can be understood through a Phase, Guarantee, and Utility (PGU) lens.
- Privacy guarantees can be analyzed from object-oriented (data/model) and pipeline-oriented (local/global/full-chain) viewpoints.
- A fourfold taxonomy (data publishing, data processing, architectural, hybrid) captures major PPML approaches and their utility trade-offs.
- Privacy-preserving data preparation often relies on anonymization or differential privacy, while crypto-based training/inference leverages HE/FE and related techniques.
- Full privacy-preserving pipelines are rare and require integration of privacy-preserving training and serving strategies.
- The paper outlines open problems and directions spanning measurement, attack/defense, efficiency, privacy-utility trade-offs, and benchmarking.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.