[Paper Review] When Machine Learning Meets Privacy: A Survey and Outlook
A comprehensive survey of privacy issues in machine learning, categorizing ML roles as protection target, protection tool, and attack tool, and outlining future research directions.
The newly emerged machine learning (e.g. deep learning) methods have become a strong driving force to revolutionize a wide range of industries, such as smart healthcare, financial technology, and surveillance systems. Meanwhile, privacy has emerged as a big concern in this machine learning-based artificial intelligence era. It is important to note that the problem of privacy preservation in the context of machine learning is quite different from that in traditional data privacy protection, as machine learning can act as both friend and foe. Currently, the work on the preservation of privacy and machine learning (ML) is still in an infancy stage, as most existing solutions only focus on privacy problems during the machine learning process. Therefore, a comprehensive study on the privacy preservation problems and machine learning is required. This paper surveys the state of the art in privacy issues and solutions for machine learning. The survey covers three categories of interactions between privacy and machine learning: (i) private machine learning, (ii) machine learning aided privacy protection, and (iii) machine learning-based privacy attack and corresponding protection schemes. The current research progress in each category is reviewed and the key challenges are identified. Finally, based on our in-depth analysis of the area of privacy and machine learning, we point out future research directions in this field.
Motivation & Objective
- Survey the state of the art on privacy in machine learning and identify key challenges.
- Categorize interactions between privacy and ML into three roles (private ML, ML-aided privacy protection, ML-based privacy attacks).
- Analyze attacks, protection schemes, and collaborative learning approaches to privacy in ML.
- Provide guidance and directions for future research in privacy-preserving ML.
Proposed method
- Classify existing works by ML roles in privacy (private ML, ML-aided privacy protection, ML-based privacy attacks).
- Review attack models and protection schemes in private ML, including model/data privacy, and various threat settings (white-box/black-box).
- Discuss encryption, obfuscation/differential privacy, and secure computation as privacy-preserving techniques for ML.
- Explain distributed and collaborative learning frameworks (federated, split learning) and their privacy implications.
- Summarize ML-aided privacy protection methods and ML-based privacy attacks to offer insights for future research.
Experimental results
Research questions
- RQ1What are the main privacy threats and roles of ML in privacy (model/data privacy, ML as protection tool, ML as attack tool) in ML systems?
- RQ2What privacy-preserving techniques and architectures (encryption, DP, SMC, federated/split learning) effectively protect ML models and data?
- RQ3How do existing works categorize and compare attacks and protections across private ML, ML-aided privacy protection, and ML-based privacy attacks?
- RQ4What are the open challenges and future directions for privacy in ML with unstructured data and complex models?
Key findings
- Privacy in ML involves three interacting roles: ML as protection target, ML as protection tool, and ML as an attack tool.
- Private ML attacks focus on training data privacy and model privacy, with threats including model extraction, feature estimation, membership inference, and model memorization.
- DP and related privacy accounting (moments accountant, Rényi DP) are central but have limitations for ML, especially with unstructured data.
- Encryption (including homomorphic encryption) and secure multi-party computation can protect data and models but introduce significant computational and communication overhead.
- Collaborative learning frameworks (federated and split learning) can reduce data exposure but raise unique privacy risks and defense needs.
- ML-aided privacy protection uses ML to identify privacy risks and adjust sharing policies, while ML-based privacy attacks exploit ML capabilities to infer sensitive data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.