[论文解读] ChatGPT-powered Conversational Drug Editing Using Retrieval and Domain Feedback
ChatDrug 使用一个由三个模块组成的框架(PDDS、ReDF 和 conversation),结合 Retrieval 与 Domain Feedback,对小分子、肽和蛋白质进行文本-guided 编辑,在 39 项任务中达到 33 项的最佳性能。
Recent advancements in conversational large language models (LLMs), such as ChatGPT, have demonstrated remarkable promise in various domains, including drug discovery. However, existing works mainly focus on investigating the capabilities of conversational LLMs on chemical reaction and retrosynthesis. While drug editing, a critical task in the drug discovery pipeline, remains largely unexplored. To bridge this gap, we propose ChatDrug, a framework to facilitate the systematic investigation of drug editing using LLMs. ChatDrug jointly leverages a prompt module, a retrieval and domain feedback (ReDF) module, and a conversation module to streamline effective drug editing. We empirically show that ChatDrug reaches the best performance on 33 out of 39 drug editing tasks, encompassing small molecules, peptides, and proteins. We further demonstrate, through 10 case studies, that ChatDrug can successfully identify the key substructures (e.g., the molecule functional groups, peptide motifs, and protein structures) for manipulation, generating diverse and valid suggestions for drug editing. Promisingly, we also show that ChatDrug can offer insightful explanations from a domain-specific perspective, enhancing interpretability and enabling informed decision-making. This research sheds light on the potential of ChatGPT and conversational LLMs for drug editing. It paves the way for a more efficient and collaborative drug discovery pipeline, contributing to the advancement of pharmaceutical research and development.
研究动机与目标
- 将 AI 辅助的药物编辑视为一个多模态的、超越传统结构中心方法的对话任务。
- 开发一个提示设计与检索增强的系统,以引导 ChatGPT 进行药物编辑。
- 展示在小分子、肽和蛋白质上广泛有效性,并具备领域感知的反馈回路。
- 展示可解释性与案例研究,突出药物编辑中的子结构与基序识别。
提出的方法
- PDDS:设计领域特定的提示,以引导 ChatGPT 在三种药物类型上实现高层次性质编辑。
- ReDF:检索结构相似的候选分子并将领域反馈注入提示中以引导生成。
- 对话模块:实现迭代轮次,其中失败的编辑触发基于检索的 refinements 与再提示。
- 使用涵盖小分子、肽和蛋白质的 39 项编辑任务的多任务基准进行评估。
- 将 ChatDrug 视为一个无参数、基于提示工程的系统,不进行学习。
实验结果
研究问题
- RQ1一个对话式的大语言模型是否可以通过领域感知的提示和基于检索的反馈,有效地引导在小分子、肽和蛋白质方面进行药物编辑?
- RQ2与零-shot 与上下文学习基线相比,检索与领域反馈循环是否能提升编辑质量与多样性?
- RQ3对话轮次与反馈阈值对编辑性能有何影响?
- RQ4ChatDrug 是否能够提供可解释、与领域相关的解释,并识别参与编辑的关键子结构或基序?
主要发现
- ChatDrug 在 39 项药物编辑任务中,在小分子、肽和蛋白质领域均达到最佳性能,为 33 项任务。
- 定性案例研究显示,ChatDrug 能识别出关键子结构、基序或蛋白质区域,这些区域负责实现期望的编辑。
- 在分子编辑方面,ChatDrug 在多项单目标与多目标任务上均优于基线,且比随机或标准 MoleculeSTM 基线有显著改进。
- 对于肽,ChatDrug 的编辑增强了结合基序,与实验的肽-主要组织相容性(peptide-MHC)结合趋势相符。
- 对于蛋白质,ChatDrug 的编辑在下游预测器和折叠工具评估中表现为二级结构含量增加(更多的螺旋或更多的片段)。
- 消融研究表明,对话式细化与领域反馈注入显著优于零-shot 与单轮检索方法,随着轮次增加(直至收敛点 C=2)性能有所提升。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。