Skip to main content
QUICK REVIEW

[论文解读] Autofocused oracles for model-based design

Clara Fannjiang, Jennifer Listgarten|arXiv (Cornell University)|Jun 14, 2020
Image Processing Techniques and Applications参考文献 51被引用 17
一句话总结

本文提出自适应聚焦(autofocusing),一种在基于模型的设计过程中动态更新回归型oracle的方法,以提升设计空间分布外区域的性能。通过将设计建模为非零和博弈,并利用高置信度预测结果迭代重训练oracle,自适应聚焦显著提高了预测准确率,并增强了对高性能候选物的发现能力。在高温超导材料与蛋白质设计基准测试中,其表现达到中位数临界温度预测提升最高50%,且有5.3%的概率超越训练数据的最大值。

ABSTRACT

Data-driven design is making headway into a number of application areas, including protein, small-molecule, and materials engineering. The design goal is to construct an object with desired properties, such as a protein that binds to a therapeutic target, or a superconducting material with a higher critical temperature than previously observed. To that end, costly experimental measurements are being replaced with calls to high-capacity regression models trained on labeled data, which can be leveraged in an in silico search for design candidates. However, the design goal necessitates moving into regions of the design space beyond where such models were trained. Therefore, one can ask: should the regression model be altered as the design algorithm explores the design space, in the absence of new data? Herein, we answer this question in the affirmative. In particular, we (i) formalize the data-driven design problem as a non-zero-sum game, (ii) develop a principled strategy for retraining the regression model as the design algorithm proceeds---what we refer to as autofocusing, and (iii) demonstrate the promise of autofocusing empirically.

研究动机与目标

  • 为解决在远离训练数据区域的oracle预测不可靠这一数据驱动设计中的常见问题。
  • 开发一种在设计过程探索新区域时,无需新增标注数据即可更新oracle模型的系统性策略。
  • 将设计过程形式化为设计算法与oracle之间的非零和博弈,以实现联合优化。
  • 通过实证验证,迭代重训练oracle(即自适应聚焦)可提升所发现设计的质量。
  • 展示自适应聚焦如何同时提升预测准确率与真实世界设计任务中高性能候选物的发现能力。

提出的方法

  • 将基于模型的设计形式化为设计算法(搜索模型)与oracle(回归模型)之间的非零和博弈,双方均旨在优化其各自目标。
  • 基于博弈论中的纳什均衡推导oracle更新策略,采用加权最大似然估计法,利用搜索模型生成的高置信度预测结果对oracle进行重训练。
  • 通过定期使用当前搜索分布中预测属性值较高的样本重训练oracle,实现自适应聚焦的实施,优先关注感兴趣区域。
  • 使用搜索模型(如Potts模型或变分自编码器)生成多样化候选设计,并引导oracle的更新过程。
  • 在多种优化算法(如RWR、CMA-ES、CEM-PI)上应用该方法,并在蛋白质设计与高温超导材料设计任务中评估性能。
  • 采用中位数与最大预测属性值、改善概率百分比(PCI)、斯皮尔曼等级相关系数及与真实测量值的均方根误差(RMSE)等指标评估结果。

实验结果

研究问题

  • RQ1迭代重训练oracle模型是否能提升在远离训练数据区域的设计性能?
  • RQ2将设计建模为非零和博弈的博弈论框架,如何为oracle更新提供系统性策略?
  • RQ3与静态或固定oracle相比,自适应聚焦是否能带来更优的高性能设计发现能力?
  • RQ4自适应聚焦在多大程度上提升了分布外区域的预测准确率与可靠性?
  • RQ5自适应聚焦在不同基于模型的优化算法与设计任务中的表现如何变化?

主要发现

  • 与非自适应聚焦基线相比,自适应聚焦使高温超导材料的中位数预测临界温度提升了30.7 K(p < 0.01)。
  • 在同一任务中,最大预测临界温度提升了18.6 K(p < 0.01),有4.3%的概率超越训练数据的最大值。
  • 在蛋白质设计任务中,自适应聚焦oracle与真实值的斯皮尔曼等级相关系数达到0.55,较基线提升0.52(p < 0.01)。
  • 在高温超导材料任务中,自适应聚焦使oracle与真实值预测之间的均方根误差(RMSE)降低了6.1 K(p < 0.01)。
  • 在CMA-ES优化方法中,自适应聚焦使中位数预测属性值提升14.1 K(p < 0.05),最大值提升15.9 K(p < 0.05)。
  • 该方法在所有测试的六种MBO方法中均持续优于非自适应聚焦基线,在18项性能指标中的15项达到统计显著提升。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。