Skip to main content
QUICK REVIEW

[论文解读] Large-scale Feature Selection of Risk Genetic Factors for Alzheimer's Disease via Distributed Group Lasso Regression

Qingyang Li, Dajiang Zhu|arXiv (Cornell University)|Apr 27, 2017
Cancer-related molecular mechanisms research参考文献 10被引用 3
一句话总结

本文提出了一种用于跨多个机构的大规模阿尔茨海默病(AD)遗传风险因子发现的分布式组套索特征选择框架(DFSF)。通过整合分布式筛选规则与稳定性选择,DFSF在ADNI的590万 SNP 数据集中,相较于ADMM实现了35至38倍的加速,同时识别出APOE、GRM8、GPC6和LOC100506272等SNP为显著的AD风险因子。

ABSTRACT

Genome-wide association studies (GWAS) have achieved great success in the genetic study of Alzheimer's disease (AD). Collaborative imaging genetics studies across different research institutions show the effectiveness of detecting genetic risk factors. However, the high dimensionality of GWAS data poses significant challenges in detecting risk SNPs for AD. Selecting relevant features is crucial in predicting the response variable. In this study, we propose a novel Distributed Feature Selection Framework (DFSF) to conduct the large-scale imaging genetics studies across multiple institutions. To speed up the learning process, we propose a family of distributed group Lasso screening rules to identify irrelevant features and remove them from the optimization. Then we select the relevant group features by performing the group Lasso feature selection process in a sequence of parameters. Finally, we employ the stability selection to rank the top risk SNPs that might help detect the early stage of AD. To the best of our knowledge, this is the first distributed feature selection model integrated with group Lasso feature selection as well as detecting the risk genetic factors across multiple research institutions system. Empirical studies are conducted on 809 subjects with 5.9 million SNPs which are distributed across several individual institutions, demonstrating the efficiency and effectiveness of the proposed method.

研究动机与目标

  • 解决多中心影像遗传学研究中大规模、隐私保护的特征选择挑战,用于阿尔茨海默病(AD)研究。
  • 克服在分布式机构间分析高维GWAS数据时面临的计算与隐私障碍。
  • 开发一种可扩展的分布式框架,将组套索与稳定性选择相结合,以识别具有生物学意义的SNP组。
  • 通过优先筛选多个研究机构中的相关遗传特征,实现在早期阶段高效检测AD风险因子。

提出的方法

  • 提出一种分布式特征选择框架(DFSF),可在不共享原始遗传数据的前提下,实现跨机构的协作分析。
  • 引入分布式组套索筛选规则(DSR和DDPP_GL),在优化前识别并移除不活跃特征。
  • 采用DBCD(分布式块坐标下降)方法,实现具有收敛性保证的高效分布式组套索优化。
  • 在一系列正则化参数上应用稳定性选择,基于选择频率对表现最佳的SNP进行排序。
  • 采用分布式本地查询模型(LQM)以保护数据隐私,同时支持跨机构联合学习。
  • 在Apache Spark上构建可扩展的处理流程,使用三个机构共30个节点的计算资源,处理来自809名ADNI受试者的590万 SNP 数据。

实验结果

研究问题

  • RQ1分布式组套索框架是否能在保护数据隐私的前提下,有效识别多个数据共享机构中与阿尔茨海默病相关的遗传风险因子?
  • RQ2分布式筛选规则的整合在高维GWAS数据中如何提升大规模组套索的效率?
  • RQ3稳定性选择在分布式组特征选择中在多大程度上提升了SNP排序的可靠性?
  • RQ4在多中心环境下,所提出的DFSF框架与最先进的分布式求解器(如ADMM)相比,在速度和准确性方面表现如何?
  • RQ5在不同脑区体积表型下,使用该方法是否能一致地识别出特定的SNP作为AD的顶级风险因子?

主要发现

  • 在三个机构处理590万 SNP 数据时,所提出的DFSF相较于ADMM求解器实现了38倍的加速。
  • 稳定性选择在海马体和内嗅皮层体积预测中均将APOE列为排名第一的SNP,证实了其与AD的已知关联性。
  • 该方法还检测到其他具有生物学意义的SNP,如GRM8、GPC6、LOC100506272和PIK3C2G,这些结果得到了先前GWAS研究的支持。
  • DDPP_GL+DBCD在SNP排序方面优于D_EDPP+F_LQM,尤其在识别标准套索方法未能捕捉到的非-APOE风险基因方面表现更优。
  • 该框架通过早期利用分布式筛选规则剔除不活跃SNP,成功缩减了特征空间,加速了收敛过程,且未损失准确性。
  • 在ADNI数据上的实证验证表明,该方法在真实世界多中心影像遗传学研究中具备良好的可扩展性与有效性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。