[论文解读] RECOVER: sequential model optimization platform for combination drug repurposing identifies novel synergistic compounds in vitro
RECOVER 是一个顺序模型优化平台,利用深度学习优先筛选可用于药物重定位的协同药物组合,将体外筛选范围减少至总搜索空间的约5%。经过三轮机器学习引导的实验,该平台识别出高度协同的药物对,重新发现了后来在临床试验中被证实有效的组合,并表明基于结构的药物嵌入能够反映药物作用的生物学机制。
For large libraries of small molecules, exhaustive combinatorial chemical screens become infeasible to perform when considering a range of disease models, assay conditions, and dose ranges. Deep learning models have achieved state of the art results in silico for the prediction of synergy scores. However, databases of drug combinations are biased towards synergistic agents and these results do not necessarily generalise out of distribution. We employ a sequential model optimization search utilising a deep learning model to quickly discover synergistic drug combinations active against a cancer cell line, requiring substantially less screening than an exhaustive evaluation. Our small scale wet lab experiments only account for evaluation of ~5% of the total search space. After only 3 rounds of ML-guided in vitro experimentation (including a calibration round), we find that the set of drug pairs queried is enriched for highly synergistic combinations; two additional rounds of ML-guided experiments were performed to ensure reproducibility of trends. Remarkably, we rediscover drug combinations later confirmed to be under study within clinical trials. Moreover, we find that drug embeddings generated using only structural information begin to reflect mechanisms of action. Prior in silico benchmarking suggests we can enrich search queries by a factor of ~5-10x for highly synergistic drug combinations by using sequential rounds of evaluation when compared to random selection, or by a factor of >3x when using a pretrained model selecting all drug combinations at a single time point.
研究动机与目标
- 为解决在癌症药物重定位中大规模药物组合库进行穷举筛选不可行的问题。
- 克服现有药物协同数据库中的数据偏差,以提升机器学习模型的泛化能力。
- 通过使用主动学习的顺序模型优化(SMO)减少实验负担,优先筛选高潜力药物对。
- 评估仅基于结构学习的药物嵌入是否能反映生物学作用机制。
- 展示从先前数据集迁移学习的潜力,以提升在新实验设置下的性能。
提出的方法
- 使用药物结构指纹和细胞系特征,通过深度学习模型预测协同得分,训练数据来自公开的药物组合数据集。
- 顺序模型优化(SMO)利用平衡探索与利用的获取函数,选择用于体外测试的药物对。
- 在每轮湿实验后迭代重新训练模型,以提高预测准确性,并聚焦于搜索空间中高协同区域。
- 通过深度集成的不确定性估计指导探索,确保选择多样化且信息丰富的查询。
- 通过在 O’Neil 数据集上预训练模型,再在 NCI-ALMANAC 数据集上微调,实现迁移学习,以提升在新设置下的性能。
- 从分子结构中学习药物嵌入,并分析其与已知作用机制的相关性。
实验结果
研究问题
- RQ1顺序模型优化流程是否能显著减少识别协同药物组合所需的体外实验数量?
- RQ2在存在偏差的药物组合数据库上训练的深度学习模型,在新实验设置下的泛化能力如何?
- RQ3基于结构的药物嵌入能否捕捉到具有生物相关性的作用机制?
- RQ4在先前数据集上预训练是否能提升在新药物组合筛选任务中的性能?
- RQ5在药物协同发现的主动学习中,基于不确定性的获取函数是否优于贪婪策略?
主要发现
- 仅约5%的总药物组合空间被实验评估,但经过三轮机器学习引导的实验后,该方法成功富集了高度协同的药物对。
- 与随机选择相比,该平台在识别顶级协同组合方面实现了5–10倍的富集,与单次时间点模型选择相比,富集度超过3倍。
- 额外两轮实验验证了观察到的协同趋势具有可重复性。
- 该方法成功重新发现了后来被证实正在临床试验中的药物组合。
- 仅从分子结构学习到的药物嵌入开始反映已知的作用机制,表明其具有潜在的生物学相关性。
- 在 O’Neil 数据集上预训练显著提升了在 NCI-ALMANAC 数据集上的性能,尽管存在分布偏移,仍证明了迁移学习的潜力。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。