[論文レビュー] Accurate Clinical Toxicity Prediction using Multi-task Deep Neural Nets and Contrastive Molecular Explanations
本稿では、Morganフィンガープrintと事前学習済みSMILES埋め込みを用いたマルチタスク深層ニューラルネットワークを提案し、in vitro、in vivo、臨床毒性を同時に予測するもので、MoleculeNetベンチマークにおいて最先端の性能を達成した。また、既知のトキシコフォアに一致するPertinent Positive(関連する正例)およびPertinent Negative(関連しない負例)特徴を用いた対比的説明を導入し、モデルの解釈性を向上させるとともに、臨床エンドポイントよりもin vitro/in vivoのトキシコフォア回復を優先する傾向を明らかにした。
Explainable ML for molecular toxicity prediction is a promising approach for efficient drug development and chemical safety. A predictive ML model of toxicity can reduce experimental cost and time while mitigating ethical concerns by significantly reducing animal and clinical testing. Herein, we use a deep learning framework for simultaneously modeling in vitro, in vivo, and clinical toxicity data. Two different molecular input representations are used: Morgan fingerprints and pre-training SMILES embeddings. A multi-task deep learning model accurately predicts toxicity for all endpoints, including clinical, as indicated by AUROC and balanced accuracy. In particular, SMILES embeddings as input to the multi-task model improved clinical toxicity predictions compared to existing models in MoleculeNet benchmark. Additionally, our multi-task approach is comprehensive in the sense that it is comparable to state-of-the-art approaches for specific endpoints in in vitro, in vivo and clinical platforms. Through both the multi-task model and transfer learning, we were able to indicate the minimal need of in vivo data for clinical toxicity predictions. To provide confidence and explain the model's predictions, we adapt a post-hoc contrastive explanation method that returns pertinent positive and pertinent negative features, which correspond well to known mutagenic and reactive toxicophores, such as unsubstituted bonded heteroatoms, aromatic amines, and Michael receptors. Furthermore, toxicophore recovery by pertinent feature analysis captures more of the in vitro (53%) and in vivo (56%), rather than of the clinical (8%), endpoints, and indeed uncovers a preference in known toxicophore data towards in vitro and in vivo experimental data. To our knowledge, this is the first contrastive explanation, using both present and absent substructures, for predictions of clinical and in vivo molecular toxicity.
研究の動機と目的
- 薬物候補の毒性による高い失敗率を鑑み、機械学習を用いて臨床毒性を正確に予測する課題に取り組む。
- in vitro、in vivo、臨床毒性エンドポイントを同時にモデル化する統合された深層学習フレームワークを構築し、単一タスクモデルの限界を克服する。
- Pertinent Positive(存在する)およびPertinent Negative(存在しない)分子部分構造を特定する対比的説明法を用いて、毒性予測に対する実用的で忠実な説明を提供する。
- 転移学習とマルチタスク学習を活用することで、高コストかつ倫理的に懸念されるin vivoデータへの依存を低減する。
- モデルの説明が既知のミュタゲン的および反応性トキシコフォアと一致することを検証し、ドラッグ開発における信頼性と実用性を高める。
提案手法
- 3つの毒性エンドポイント(Tox21(in vitro)、RTECS(in vivo)、ClinTox(臨床))を対象に、共有層とタスク固有層を用いたマルチタスク深層ニューラルネットワーク(MT-DNN)を訓練した。
- 2種類の分子入力表現を用いた:Morganフィンガープリントと事前学習済みSMILES埋め込みで、後者が臨床毒性予測において優れた性能を示した。
- 転移学習を適用し、Tox21およびRTECSでの事前学習後に、1エポック(1)の微調整をClinToxタスクで実施した。
- 対比的説明法(CEM)を改変し、L1およびL2正則化と自己符号化器ベースの再構成損失を用いた最適化問題を解くことで、Pertinent Positive(PP)およびPertinent Negative(PN)特徴を生成した。
- 最適化には、予測を変化させる最小限で必要な摂動を求めるために、投影型高速イテレーティブシェイピングスラッシュスレッディングアルゴリズム(FISTA)を用いた。
- 自己符号化器を用いて摂動された分子が化学的に妥当であることを保証し、再構成損失をγ重み付き正則化で強制した。
実験結果
リサーチクエスチョン
- RQ1共有およびタスク固有の表現を用いて、in vitro、in vivo、臨床プラットフォームの毒性を同時にかつ正確に予測できるマルチタスク深層学習モデルは構築可能か?
- RQ2従来の分子フィンガープリントと比較して、事前学習済みSMILES埋め込みを用いることで、臨床毒性予測性能が向上するか?
- RQ3転移学習を用いることで、in vivoデータの必要性をどの程度低減できるか?
- RQ4対比的説明(Pertinent PositiveおよびNegative)は、置換されていないヘテロ原子、アロマティックアミン、Michael受容体などの既知のトキシコフォアを特定できるか?
- RQ5なぜモデルの説明性能はin vitroおよびin vivoのトキシコフォアよりも臨床のものと相関が弱いか?
主な発見
- 事前学習済みSMILES埋め込みを用いたマルチタスクモデルは、MoleculeNetベンチマークにおいて臨床毒性予測で最先端の性能を達成した。
- 全プラットフォームで高い予測精度を示し、AUCおよびバランス精度指標が優れた一般化性能を示した。
- Pertinent PositiveおよびNegative特徴によるトキシコフォア回復では、既知のin vitroトキシコフォアの53%とin vivoトキシコフォアの56%を捉えたが、臨床トキシコフォアはたったの8%にとどまった。
- トキシコフォア回復率の差は、in vitroおよびin vivoの実験データに偏った既知のトキシコフォアデータベースのバイアスを示唆している。
- 対比的説明は、アロマティックアミン、置換されていないヘテロ原子、Michael受容体などの生物学的に関連する部分構造が毒性予測の主な寄与要因であることを明確にした。
- 転移学習の活用により、最小限のin vivoデータで効果的な臨床毒性予測が可能となり、モデルのデータ効率性と動物実験の削減可能性が示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。