Skip to main content
QUICK REVIEW

[論文レビュー] Models Matter: The Impact of Single-Step Retrosynthesis on Synthesis Planning

Paula Torren-Peraire, Alan Kai Hassen|arXiv (Cornell University)|Aug 10, 2023
Computational Drug Discovery Methods被引用数 6
ひとこと要約

本研究では、最先端の1段階後向き合成モデルを複数段階の合成計画に統合し、モデル選択がルート探索成功に顕著な影響を及ぼすことを示した。成功率は最大28%向上した。また、1段階性能と実際の計画効率との間には深刻な乖離が存在することが明らかになった。本研究はUSPTO-50kベンチマークの妥当性に疑問を呈し、今後のモデル開発を導くために、完全な合成計画パイプライン内での評価を提唱する。

ABSTRACT

Retrosynthesis consists of breaking down a chemical compound recursively step-by-step into molecular precursors until a set of commercially available molecules is found with the goal to provide a synthesis route. Its two primary research directions, single-step retrosynthesis prediction, which models the chemical reaction logic, and multi-step synthesis planning, which tries to find the correct sequence of reactions, are inherently intertwined. Still, this connection is not reflected in contemporary research. In this work, we combine these two major research directions by applying multiple single-step retrosynthesis models within multi-step synthesis planning and analyzing their impact using public and proprietary reaction data. We find a disconnection between high single-step performance and potential route-finding success, suggesting that single-step models must be evaluated within synthesis planning in the future. Furthermore, we show that the commonly used single-step retrosynthesis benchmark dataset USPTO-50k is insufficient as this evaluation task does not represent model performance and scalability on larger and more diverse datasets. For multi-step synthesis planning, we show that the choice of the single-step model can improve the overall success rate of synthesis planning by up to +28% compared to the commonly used baseline model. Finally, we show that each single-step model finds unique synthesis routes, and differs in aspects such as route-finding success, the number of found synthesis routes, and chemical validity, making the combination of single-step retrosynthesis prediction and multi-step synthesis planning a crucial aspect when developing future methods.

研究の動機と目的

  • 1段階後向き合成モデルの選択が複数段階の合成計画性能に与える影響を調査すること。
  • 高い性能を示す1段階モデルが、現実世界のスケーラブルな合成計画タスクに一般化するかを評価すること。
  • USPTO-50kベンチマークが現実世界のスケーラビリティと性能の転送性を適切に反映していないことの限界を評価すること。
  • 合成計画におけるモデル性能、推論速度、ルート多様性の間のトレードオフを特定すること。
  • 代表的なサブサンプルを用いた複数段階計画文脈における1段階モデルのベンチマークフレームワークを提供すること。

提案手法

  • 本研究では、NeuralSym, LocalRetro, MHNreact, Chemformer、およびベースラインのAZFという、複数の最先端の1段階後向き合成モデルを、複数段階の合成計画フレームワークに統合した。
  • 合成計画成功確率、ルート多様性、化学的妥当性を指標として、公開データセット(USPTO-50k, Caspyrus10k)および独自の反応データセットを用いてモデル性能を評価した。
  • Caspyrus10kデータセットに対してサブサンプリング戦略を採用し、1000化合物を代表的サブセットとしてテストすることで、低分散かつ高速なベンチマークを可能にした。
  • モデル間でのルート探索成功確率、生成可能な有効ルート数、推論時間、基準ルートからの乖離を分析した。
  • 公平な比較を確保するため、ビームサーチと化学的妥当性のフィルタリングを組み合わせた標準化された合成計画パイプラインを採用した。
  • SMILES表現を用い、テンプレートベースとテンプレートフリーの両アプローチを評価することで、性能と速度のトレードオフを評価した。
Figure 1: Evaluation Framework for single-step models (AiZynthFinder (AZF), LocalRetro, Chemformer, and MHNreact), trained on different public (USPTO-50k, USPTO-PaRoutes-1M) and proprietary (AZ-1M, AZ-18M) datasets in synthesis planning on Caspyrus10k and PaRoutes.
Figure 1: Evaluation Framework for single-step models (AiZynthFinder (AZF), LocalRetro, Chemformer, and MHNreact), trained on different public (USPTO-50k, USPTO-PaRoutes-1M) and proprietary (AZ-1M, AZ-18M) datasets in synthesis planning on Caspyrus10k and PaRoutes.

実験結果

リサーチクエスチョン

  • RQ1USPTO-50kのような小規模なベンチマークで高い1段階後向き合成性能を示すモデルは、複数段階の合成計画においても成功するか?
  • RQ2複数段階計画において、テンプレートベースとテンプレートフリーの1段階モデルは、ルート探索成功、速度、多様性の観点でどのように比較できるか?
  • RQ3USPTO-50kデータセットは、より大規模かつ多様な反応データセット(USPTO-PaRoutes-1M や Caspyrus10k)における現実世界のスケーラビリティと性能の転送性を適切に反映しているか?
  • RQ41段階モデルの選択が、全体の合成計画成功確率を顕著に向上させることができるか。その向上率はどの程度か?
  • RQ5異なる1段階モデルは、独自の有効な合成ルートを生成するか。また、主な性能指標(成功確率、有効ルート数、推論時間、化学的妥当性)においてどのように異なるか?

主な発見

  • 同じ反応データで学習されたとしても、1段階後向き合成モデルの選択により、複数段階の合成計画成功確率を最大28%向上させることができる。
  • USPTO-50kでの高い1段階性能と、複数段階計画における成功したルート探索との間に明確な相関はなく、評価フレームワークに深刻な乖離が存在することが示された。
  • USPTO-50kベンチマークは、現実世界のスケーラビリティと性能の転送性を評価するには不十分であり、USPTO-PaRoutes-1M や Caspyrus10k などの大規模・多様なデータセットでは、性能順位やモデルの挙動が一般化しないことが判明した。
  • テンプレートフリーのモデル(例:Chemformer)は、大規模かつ多様なデータに対して優れた1段階性能を示すが、複数段階計画では200倍も遅くなる。一方、テンプレートベースのモデル(例:LocalRetro)は、速度、成功確率、ルート多様性のバランスに優れている。
  • 各1段階モデルは、成功確率、有効ルート数、推論時間、化学的妥当性の観点で異なる合成ルートを生成する。これは、計画におけるモデル選択の重要性を強調している。
  • Caspyrus10kの1000化合物の代表的サブセットは、全データセットの結果を安定的に近似でき、成功確率の標準偏差が0.05未満という低分散で、高速なベンチマークが可能である。
Figure 2: Single-step Retrosynthesis Prediction Performance in terms of top-n accuracy for AZF, LocalRetro, Chemformer, and MHNreact on different datasets (USPTO-50k, USPTO-PaRoutes-1M, AZ-1M, AZ-18M) (see Supplementary Table S1 ).
Figure 2: Single-step Retrosynthesis Prediction Performance in terms of top-n accuracy for AZF, LocalRetro, Chemformer, and MHNreact on different datasets (USPTO-50k, USPTO-PaRoutes-1M, AZ-1M, AZ-18M) (see Supplementary Table S1 ).

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。