[Paper Review] Assessing the risk of re-identification arising from an attack on anonymised data
This paper proposes an analytical framework to estimate the risk of re-identification in k-anonymized electronic health record (EHR) datasets under targeted attacks. Using recursive probability models and Monte Carlo simulations, it quantifies how re-identification likelihood increases with the proportion of data obtained by an adversary and decreases with larger equivalence class sizes, enabling systematic k-anonymization parameterization based on desired risk thresholds.
Objective: The use of routinely-acquired medical data for research purposes requires the protection of patient confidentiality via data anonymisation. The objective of this work is to calculate the risk of re-identification arising from a malicious attack to an anonymised dataset, as described below. Methods: We first present an analytical means of estimating the probability of re-identification of a single patient in a k-anonymised dataset of Electronic Health Record (EHR) data. Second, we generalize this solution to obtain the probability of multiple patients being re-identified. We provide synthetic validation via Monte Carlo simulations to illustrate the accuracy of the estimates obtained. Results: The proposed analytical framework for risk estimation provides re-identification probabilities that are in agreement with those provided by simulation in a number of scenarios. Our work is limited by conservative assumptions which inflate the re-identification probability. Discussion: Our estimates show that the re-identification probability increases with the proportion of the dataset maliciously obtained and that it has an inverse relationship with the equivalence class size. Our recursive approach extends the applicability domain to the general case of a multi-patient re-identification attack in an arbitrary k-anonymisation scheme. Conclusion: We prescribe a systematic way to parametrize the k-anonymisation process based on a pre-determined re-identification probability. We observed that the benefits of a reduced re-identification risk that come with increasing k-size may not be worth the reduction in data granularity when one is considering benchmarking the re-identification probability on the size of the portion of the dataset maliciously obtained by the adversary.
Motivation & Objective
- To quantify the risk of re-identification in anonymized medical datasets under a malicious data breach scenario.
- To develop an analytical method for estimating the probability of re-identification for a single patient in a k-anonymized EHR dataset.
- To generalize the single-patient model to estimate the risk of re-identifying multiple patients simultaneously.
- To validate the analytical estimates using synthetic Monte Carlo simulations.
- To guide the systematic selection of k-anonymization parameters based on a pre-defined re-identification risk threshold.
Proposed method
- Derives an analytical expression for the probability of re-identification of a single patient in a k-anonymized dataset using combinatorial probability.
- Generalizes the single-patient model into a recursive formulation to compute the probability of re-identifying multiple patients in arbitrary k-anonymization schemes.
- Employs Monte Carlo simulations on synthetic EHR datasets to validate the accuracy of the analytical risk estimates across various attack scenarios.
- Assesses the impact of key variables such as the proportion of dataset obtained by an adversary and the size of equivalence classes on re-identification risk.
- Uses conservative assumptions in the model to ensure upper-bound estimation of re-identification probability, enhancing safety in risk assessment.
- Proposes a systematic method to parametrize k-anonymization based on a target re-identification risk level.
Experimental results
Research questions
- RQ1What is the analytical probability of re-identifying a single patient in a k-anonymized EHR dataset under a targeted attack?
- RQ2How does the probability of re-identification scale when multiple patients are targeted simultaneously?
- RQ3How accurate are the analytical risk estimates compared to empirical simulation results across varying attack conditions?
- RQ4How does the proportion of data obtained by an adversary affect the re-identification risk in k-anonymized datasets?
- RQ5What is the trade-off between reducing re-identification risk through increased k and the resulting loss of data utility?
Key findings
- The analytical model's re-identification probability estimates closely align with those obtained from Monte Carlo simulations across multiple test scenarios.
- Re-identification risk increases significantly with the proportion of the dataset maliciously obtained by an adversary.
- There is an inverse relationship between re-identification risk and the size of the equivalence class in k-anonymization.
- The recursive formulation successfully extends the model to handle multi-patient re-identification attacks in arbitrary k-anonymization schemes.
- The benefits of reducing re-identification risk by increasing k may be outweighed by the loss of data granularity when considering partial dataset breaches.
- The proposed framework enables systematic configuration of k-anonymization parameters based on a predefined acceptable risk threshold.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.