Skip to main content
QUICK REVIEW

[论文解读] Semi-supervised Triply Robust Inductive Transfer Learning

Tianxi Cai, Mengyan Li|arXiv (Cornell University)|Sep 12, 2022
Cancer-related molecular mechanisms research被引用 4
一句话总结

本文提出STRIFLE,一种半监督、三重稳健的归纳迁移学习方法,在高维协变量偏移条件下,通过整合来自源人群的标注数据与目标人群的未标注数据,提升代表性不足人群的预测准确性。通过结合密度比模型与插补模型,STRIFLE在两个正态模型均误设时仍能保持稳健性,其性能保证不低于仅使用目标数据的半监督学习,且误差可忽略不计。

ABSTRACT

In this work, we propose a Semi-supervised Triply Robust Inductive transFer LEarning (STRIFLE) approach, which integrates heterogeneous data from a label-rich source population and a label-scarce target population and utilizes a large amount of unlabeled data simultaneously to improve the learning accuracy in the target population. Specifically, we consider a high dimensional covariate shift setting and employ two nuisance models, a density ratio model and an imputation model, to combine transfer learning and surrogate-assisted semi-supervised learning strategies effectively and achieve triple robustness. While the STRIFLE approach assumes the target and source populations to share the same conditional distribution of outcome Y given both the surrogate features S and predictors X, it allows the true underlying model of Y|X to differ between the two populations due to the potential covariate shift in S and X. Different from double robustness, even if both nuisance models are misspecified or the distribution of Y|(S, X) is not the same between the two populations, the triply robust STRIFLE estimator can still partially use the source population when the shifted source population and the target population share enough similarities. Moreover, it is guaranteed to be no worse than the target-only surrogate-assisted semi-supervised estimator with an additional error term from transferability detection. These desirable properties of our estimator are established theoretically and verified in finite samples via extensive simulation studies. We utilize the STRIFLE estimator to train a Type II diabetes polygenic risk prediction model for the African American target population by transferring knowledge from electronic health records linked genomic data observed in a larger European source population.

研究动机与目标

  • 解决精准医学中因数据异质性和协变量偏移导致代表性不足人群性能不佳的挑战。
  • 通过利用大量未标注数据的代理辅助半监督学习,克服临床结局标注数据稀缺的问题。
  • 开发一种迁移学习框架,即使在两个正态模型(密度比模型与插补模型)均误设时仍能保持准确性。
  • 确保估计量在模型误设情况下性能不劣于仅使用目标数据的半监督学习。
  • 通过从更大、标注丰富的源人群迁移知识,实现在目标人群中高维风险预测。

提出的方法

  • 在高维协变量偏移设定下,整合来自标注丰富源人群与标注稀缺目标人群的异质性数据。
  • 采用两个正态模型:密度比模型用于校正源与目标人群之间的协变量偏移,插补模型用于利用代理特征预测结果。
  • 通过双重稳健估计方程结合迁移学习与代理辅助半监督学习,实现三重稳健性。
  • 确保若两个正态模型中至少有一个正确设定,或结果在给定预测变量与代理特征下的条件分布在不同人群中相同,则估计量保持一致。
  • 采用惩罚性估计方程方法处理高维预测变量与代理特征,实现在高维设定下的变量选择与估计。
  • 在足够人群相似性条件下,即使两个正态模型均误设,最终估计量也保证不劣于仅使用目标数据的半监督估计量。
Figure 1: The estimates of ${\bm{\beta}}_{0}$ for top signals with absolute magnitude above 0.1 from SUP, CS, SUPTrans, SAS, and STIFLE with SUPSource as a reference.
Figure 1: The estimates of ${\bm{\beta}}_{0}$ for top signals with absolute magnitude above 0.1 from SUP, CS, SUPTrans, SAS, and STIFLE with SUPSource as a reference.

实验结果

研究问题

  • RQ1当密度比模型与插补模型均误设时,迁移学习方法是否仍能在目标人群中保持性能?
  • RQ2在模型误设条件下,所提出的STRIFLE估计量与仅使用目标数据的半监督学习相比,估计精度如何?
  • RQ3来自大规模、标注丰富的源人群的知识,能在多大程度上提升在小样本、标注稀缺且存在协变量偏移的目标人群中的预测准确性?
  • RQ4STRIFLE的三重稳健性是否能确保在给定预测变量与代理特征下结果的条件分布在源与目标人群中不同时仍保持可靠性能?
  • RQ5当预测变量与代理特征数量超过标注观测数时,STRIFLE在高维设定下的有效性如何?

主要发现

  • STRIFLE实现了三重稳健性:只要两个正态模型(密度比或插补)中至少一个设定正确,估计量即保持一致。
  • 即使两个正态模型均误设,且结果在给定预测变量与代理特征下的条件分布在人群中不同,STRIFLE的性能仍不劣于仅使用目标数据的半监督估计量,且误差可忽略不计。
  • 在模拟实验中,STRIFLE通过从更大的欧洲血统源人群迁移知识,显著提升了非裔美国人人群中2型糖尿病多基因风险评分的预测准确性。
  • 该方法成功识别出关键遗传标记rs7903146、rs174546与rs1559474作为非裔美国人人群中的显著预测因子,其效应估计与已知生物学关联一致。
  • 在模拟研究与真实世界应用中,STRIFLE估计量在性能上均优于基线方法,尤其在高维与协变量偏移条件下表现更优。
  • 该方法的稳健性在多个模拟场景中得到验证,确认其在具有异质性与有限数据的实际精准医学场景中的可靠性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。