[论文解读] Integrated machine learning pipeline for aberrant biomarker enrichment (i-mAB): characterizing clusters of differentiation within a compendium of systemic lupus erythematosus patients
本研究提出i-mAB,一种集成机器学习流程,可识别与系统性红斑狼疮(SLE)中异常CD生物标志物相关的相互依赖基因表达特征。通过基于relief的算法和正则化逻辑回归,i-mAB识别出681个预测特征(平衡准确率 = 0.822),并在独立数据集中验证了新型关联——如BCL7A(padj = 1.69e-9)和STRBP(padj = 4.63e-8)与CD22的关联——实现了治疗再定位的精准生物标志物发现。
Clusters of differentiation (CD) are cell surface biomarkers that denote key biological differences between cell types and disease state. CD-targeting therapeutic monoclonal antibodies (mAB) afford rich trans-disease repositioning opportunities. Within a compendium of systemic lupus erythematous (SLE) patients, we applied the Integrated machine learning pipeline for aberrant biomarker enrichment (i-mAB) to profile de novo gene expression features affecting CD20, CD22 and CD30 gene aberrance. First, a novel relief-based algorithm identified interdependent features(p=681) predicting treatment-na\\"ive SLE patients (balanced accuracy=0.822). We then compiled CD-associated expression profiles using regularized logistic regression and pathway enrichment analyses. On an independent general cell line model system data, we replicated associations (in silico) of BCL7A(padj=1.69e-9) and STRBP(padj=4.63e-8) with CD22; NCOA2(padj=7.00e-4), ATN1(padj=1.71e-2), and HOXC4(padj=3.34e-2) with CD30; and PHOSPHO1, a phosphatase linked to bone mineralization, with both CD22(padj=4.37e-2) and CD30(padj=7.40e-3). Utilizing carefully aggregated secondary data and leveraging a priori hypotheses, i-mAB fostered robust biomarker profiling among interdependent biological features.
研究动机与目标
- 识别治疗初治SLE患者中与异常CD生物标志物相关的相互依赖基因表达特征。
- 开发一种系统性流程,通过整合机器学习与通路富集分析,描绘CD相关表达谱。
- 通过在SLE中识别新型生物标志物-疾病关联,实现跨疾病治疗再定位。
- 在独立的通用细胞系模型系统中验证预测的关联。
提出的方法
- 应用一种新型基于relief的算法,识别出在SLE患者中预测CD20、CD22和CD30异常的681个相互依赖基因表达特征。
- 使用正则化逻辑回归,从识别出的特征中整合CD相关表达谱。
- 进行通路富集分析,以阐释所识别基因集的生物学意义。
- 利用独立的通用细胞系模型系统对预测关联进行计算机模拟验证。
- 通过差异表达分析得到的校正p值(padj)评估基因-CD关联的统计显著性。
- 利用事前假设和经过仔细整合的次级数据,增强生物标志物分析的稳健性。
实验结果
研究问题
- RQ1在治疗初治SLE患者中,哪些相互依赖的基因表达特征可预测CD20、CD22和CD30的异常表达?
- RQ2从这些预测特征中推导出的关键CD相关表达谱是什么?
- RQ3预测的基因-CD关联是否可在独立的通用细胞系模型系统中复现?
- RQ4哪些新基因与CD22、CD30或两者均表现出统计学上显著的关联?
- RQ5通路富集分析如何支持所识别基因集的生物学相关性?
主要发现
- 基于relief的算法在治疗初治SLE患者中识别出681个相互依赖特征,预测CD异常的平衡准确率为0.822。
- BCL7A(padj = 1.69e-9)和STRBP(padj = 4.63e-8)在独立验证数据集中与CD22显著相关。
- NCOA2(padj = 7.00e-4)、ATN1(padj = 1.71e-2)和HOXC4(padj = 3.34e-2)与CD30表现出显著关联。
- PHOSPHO1被确定为CD22(padj = 4.37e-2)和CD30(padj = 7.40e-3)的共享生物标志物。
- 通路富集分析支持了所识别基因集的生物学相关性,将其与疾病相关生物过程相联系。
- i-mAB流程在不同数据集中成功复现了关键关联,证明了其在生物标志物分析中的稳健性与可重复性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。