Skip to main content
QUICK REVIEW

[论文解读] Covariate Microaggregation for Logistic Regression: An Application for Analysis of Confidential Data

Paramita Saha‐Chaudhuri|arXiv (Cornell University)|Mar 17, 2016
Privacy-Preserving Technologies in Data参考文献 22被引用 3
一句话总结

本文提出了一种名为Pooled Logistic Regression(PoLoR)的隐私保护方法,用于在分布式健康数据网络中分析二元疾病结局。通过在共享聚合层面数据给中央分析中心之前,将协变量微聚合成组,PoLoR能够在不暴露个体层面信息的情况下,实现逻辑回归参数(如优势比)的一致估计,同时保持与标准统计软件的兼容性,并支持通过似然比检验进行模型选择。

ABSTRACT

In the recent past, electronic health records and distributed data networks emerged as a viable resource for medical and scientific research. As the use of confidential patient information from such sources become more common, maintaining privacy of patients is of utmost importance. For a binary disease outcome of interest, we show that the techniques of microaggregation (equivalent to specimen pooling) and \underline{Po}oled \underline{Lo}gistic \underline{R}egression (PoLoR) could be applied for analysis of large and/or distributed data while respecting patient privacy. PoLoR is exactly the same as standard logistic regression, but instead of using individual covariate level, the analysis uses microaggregated covariate level when microaggregation is conditional on the outcome status. Aggregate levels of covariates can be passed from the nodes of the network to the analysis center without revealing individual-level microdata and can be used very easily with standard softwares for estimation of disease odds ratio associated with a set of categorical or continuous covariates. Microaggregation of covariates allows for consistent estimation of the parameters of logistic regression model that can include confounders and transformation of exposure. Additionally, since the microdata can be accessed within nodes, effect modifiers can be accommodated and consistently estimated. For analysis of confidential health data, covariate microaggregation for logistic regression will provide a practical and straightforward alternative to more complicated existing options.

研究动机与目标

  • 解决在分布式健康数据网络中分析机密患者数据的同时保护患者隐私的挑战。
  • 开发一种实用且计算高效的替代方法,以替代复杂隐私保护的逻辑回归方法,适用于二元结局。
  • 仅使用协变量的聚合水平数据,实现优势比和模型参数的估计,避免共享个体层面数据。
  • 在数据隐私限制下,支持使用熟悉的统计工具进行模型选择和标准误估计。
  • 证明协变量的微聚合在大数据环境中可保持参数一致性并降低数据维度。

提出的方法

  • 在每个数据节点对协变量进行微聚合,将个体分组为大小为g的组,以生成聚合层面的协变量值。
  • 将来自每个节点内部个体层面数据的聚合协变量数据共享给中央分析中心,而非个体记录。
  • 在聚合层面协变量上拟合标准逻辑回归模型,等价于PoLoR,可产生一致的参数估计。
  • 该方法支持连续和分类协变量,包括变换和混杂因素调整。
  • 通过标准似然比检验和稳健标准误进行模型选择和推断,确保结果的有效性。
  • 该方法具有可扩展性,通过确保个体层面数据不离开节点,仅交换聚合摘要,从而实现隐私保护。

实验结果

研究问题

  • RQ1协变量的微聚合是否能在不暴露个体层面数据的情况下,实现在分布式数据网络中对逻辑回归参数的一致估计?
  • RQ2组大小g的选择如何影响PoLoR中优势比估计的偏差和精度?
  • RQ3PoLoR是否能够支持使用传统统计软件和方法进行模型选择和标准误估计?
  • RQ4PoLoR是否适用于连续和分类协变量,包括变换后的暴露变量和混杂因素?
  • RQ5PoLoR估计量的渐近性质是什么?其表现如何依赖于组数而非个体数?

主要发现

  • 当使用微聚合协变量水平而非个体层面数据时,PoLoR可产生逻辑回归系数和优势比的一致估计。
  • 在结肠癌数据集中,PoLoR在组大小g = 3和g = 4时,得到的log OR估计值与标准逻辑回归非常接近,仅在标准误上存在微小差异。
  • 该方法保持了模型的有效性,支持使用传统软件进行似然比检验和标准误估计。
  • 组大小在5至20之间可实现隐私保护与估计偏差之间的合理平衡,当g > 40时,由于向均值回归,偏差会增加。
  • PoLoR具有可扩展性,由于降低了数据维度并支持单次遍历估计,适用于大数据和资源受限的环境。
  • 该方法不适用于重复更新;新增记录需要重新聚合并从头开始重新分析。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。