[论文解读] Best practices for constructing, preparing, and evaluating protein-ligand binding affinity benchmarks
本文确立了使用自由能计算构建、准备和评估蛋白质-配体结合亲和力基准的标准化最佳实践。它引入了一个经过筛选、版本化的基准数据集(protein-ligand-benchmark)以及一个开源工具包(openff-arsenic),以确保计算方法评估的高质量、可重现性和统计稳健性,广泛适用于杂化自由能方法以及药物发现中新兴的机器学习方法。
Free energy calculations are rapidly becoming indispensable in structure-enabled drug discovery programs. As new methods, force fields, and implementations are developed, assessing their expected accuracy on real-world systems (benchmarking) becomes critical to provide users with an assessment of the accuracy expected when these methods are applied within their domain of applicability, and developers with a way to assess the expected impact of new methodologies. These assessments require construction of a benchmark - a set of well-prepared, high quality systems with corresponding experimental measurements designed to ensure the resulting calculations provide a realistic assessment of expected performance when these methods are deployed within their domains of applicability. To date, the community has not yet adopted a common standardized benchmark, and existing benchmark reports suffer from a myriad of issues, including poor data quality, limited statistical power, and statistically deficient analyses, all of which can conspire to produce benchmarks that are poorly predictive of real-world performance. Here, we address these issues by presenting guidelines for (1) curating experimental data to develop meaningful benchmark sets, (2) preparing benchmark inputs according to best practices to facilitate widespread adoption, and (3) analysis of the resulting predictions to enable statistically meaningful comparisons among methods and force fields.
研究动机与目标
- 解决蛋白质-配体结合亲和力预测缺乏标准化、高质量基准的问题。
- 通过定义数据筛选、结构准备和统计分析的最佳实践,提高计算药物发现中基准测试的可靠性与可重现性。
- 提供一个社区可采用、带版本控制的基准数据集和开源工具包,以实现对自由能方法的一致性评估。
- 通过标准化的数据收集与筛选协议,支持基准数据集的系统性改进。
提出的方法
- 从高质量的实验数据中筛选出标准化、带版本控制的基准数据集(protein-ligand-benchmark),包含结构信息与生物活性数据。
- 建立实验数据筛选的最佳实践,包括最低数据质量标准,如Iridium MT/HT分类和生物物理实验来源。
- 提供详细的结构准备协议,包括正确的质子化状态、互变异构态和电荷状态,并经专家审核验证。
- 引入统计评估框架,包括自助法置信区间以及适当的指标(如RMSE和MUE)。
- 开发开源工具包openff-arsenic,以自动化并标准化不同方法与力场之间自由能计算的评估。
- 推荐一致的可视化与报告实践,例如使用相同单位和比例绘制计算值与实验值的亲和力对比图。
实验结果
研究问题
- RQ1定义高质量、可靠的蛋白质-配体结合亲和力预测基准数据集的标准是什么?
- RQ2如何对实验数据进行筛选,以最小化偏差并确保基准测试的统计稳健性?
- RQ3何种结构准备协议可确保自由能计算中的一致性与准确性?
- RQ4应如何对基准结果进行统计分析,以实现方法与力场之间有效、可重现的比较?
- RQ5哪些开放、带版本控制且由社区维护的资源可实现计算药物发现领域基准测试的标准化?
主要发现
- 每个靶点至少包含16个配体(理想情况下为25个),且自由能窗口至少为3.0 kcal/mol(理想情况下大于5.0 kcal/mol),以确保足够的动态范围以实现有意义的评估。
- 高质量的实验数据必须来源于生物物理实验,并分类为Iridium MT或HT,以确保结果的可靠性和可重现性。
- 结构准备必须包括对质子化状态、互变异构态和电荷状态的专家验证,以防止自由能预测中的系统性误差。
- 统计分析必须对所有性能指标包含自助法置信区间,以避免对结果的过度解读。
- 开源工具包openff-arsenic可实现不同方法与力场之间自由能计算的标准化、可重现的评估。
- 所提出的框架既支持回顾性基准测试,也支持前瞻性挑战,从而提升基准测试在实际应用中的预测能力。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。