[论文解读] Assessing the risk of re-identification arising from an attack on anonymised data
本文提出了一种分析框架,用于在针对性攻击下估算 k-匿名化电子健康记录(EHR)数据集中的重新识别风险。通过递归概率模型和蒙特卡洛模拟,量化了攻击者获取的数据比例越高,重新识别可能性越大,而等价类规模越大,重新识别可能性越低,从而实现基于预定风险阈值的系统性 k-匿名化参数配置。
Objective: The use of routinely-acquired medical data for research purposes requires the protection of patient confidentiality via data anonymisation. The objective of this work is to calculate the risk of re-identification arising from a malicious attack to an anonymised dataset, as described below. Methods: We first present an analytical means of estimating the probability of re-identification of a single patient in a k-anonymised dataset of Electronic Health Record (EHR) data. Second, we generalize this solution to obtain the probability of multiple patients being re-identified. We provide synthetic validation via Monte Carlo simulations to illustrate the accuracy of the estimates obtained. Results: The proposed analytical framework for risk estimation provides re-identification probabilities that are in agreement with those provided by simulation in a number of scenarios. Our work is limited by conservative assumptions which inflate the re-identification probability. Discussion: Our estimates show that the re-identification probability increases with the proportion of the dataset maliciously obtained and that it has an inverse relationship with the equivalence class size. Our recursive approach extends the applicability domain to the general case of a multi-patient re-identification attack in an arbitrary k-anonymisation scheme. Conclusion: We prescribe a systematic way to parametrize the k-anonymisation process based on a pre-determined re-identification probability. We observed that the benefits of a reduced re-identification risk that come with increasing k-size may not be worth the reduction in data granularity when one is considering benchmarking the re-identification probability on the size of the portion of the dataset maliciously obtained by the adversary.
研究动机与目标
- 量化在恶意数据泄露场景下,对匿名化医疗数据集的重新识别风险。
- 开发一种分析方法,用于估算在 k-匿名化 EHR 数据集中单个患者被重新识别的概率。
- 将单个患者模型推广至多患者同时重新识别的风险估计。
- 通过合成蒙特卡洛模拟验证分析估计的准确性。
- 基于预设的重新识别风险阈值,指导系统性选择 k-匿名化参数。
提出的方法
- 使用组合概率推导出 k-匿名化数据集中单个患者被重新识别的概率的分析表达式。
- 将单个患者模型推广为递归公式,以计算在任意 k-匿名化方案下同时重新识别多个患者的概率。
- 在合成 EHR 数据集上进行蒙特卡洛模拟,以验证在各种攻击场景下分析风险估计的准确性。
- 评估关键变量(如攻击者获取的数据比例和等价类大小)对重新识别风险的影响。
- 在模型中采用保守假设,以确保重新识别概率的上界估计,从而增强风险评估的安全性。
- 提出一种系统性方法,基于目标重新识别风险水平对 k-匿名化参数进行参数化。
实验结果
研究问题
- RQ1在针对性攻击下,k-匿名化 EHR 数据集中单个患者被重新识别的分析概率是多少?
- RQ2当多个患者被同时针对时,重新识别概率如何变化?
- RQ3在不同攻击条件下,分析风险估计与经验模拟结果的准确性如何比较?
- RQ4攻击者获取的数据比例如何影响 k-匿名化数据集中的重新识别风险?
- RQ5通过提高 k 来降低重新识别风险与由此导致的数据可用性损失之间存在何种权衡?
主要发现
- 在多个测试场景中,分析模型的重新识别概率估计与蒙特卡洛模拟结果高度一致。
- 随着攻击者恶意获取的数据比例增加,重新识别风险显著上升。
- 重新识别风险与 k-匿名化中等价类的大小呈反比关系。
- 递归公式成功将模型扩展至处理任意 k-匿名化方案下的多患者重新识别攻击。
- 在考虑部分数据集泄露时,通过提高 k 来降低重新识别风险的收益,可能被数据粒度损失所抵消。
- 所提出的框架可基于预设可接受的风险阈值,实现 k-匿名化参数的系统性配置。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。