Skip to main content
QUICK REVIEW

[论文解读] Discovery of Self-Assembling $\pi$-Conjugated Peptides by Active Learning-Directed Coarse-Grained Molecular Simulation

Kirill Shmilovich, Rachael A. Mansbach|arXiv (Cornell University)|Jan 27, 2020
Supramolecular Self-Assembly in Materials参考文献 92被引用 4
一句话总结

本研究开发了一种主动学习框架,整合了粗粒度分子动力学、变分自编码器和贝叶斯优化,以高效识别DXXX-OPV3-XXXD家族中能自组装成高度堆叠、准一维纳米聚集体的高性能π-共轭肽。通过仅模拟8,000种可能三肽序列中的2.3%(即186个序列),该方法成功识别出优异的自组装体,并揭示了有利于小至中等疏水性残基以及特定甲硫氨酸位置靠近π-核心的设计规律。

ABSTRACT

Electronically-active organic molecules have demonstrated great promise as novel soft materials for energy harvesting and transport. Self-assembled nanoaggregates formed from $\pi$-conjugated oligopeptides composed of an aromatic core flanked by oligopeptide wings offer emergent optoelectronic properties within a water soluble and biocompatible substrate. Nanoaggregate properties can be controlled by tuning core chemistry and peptide composition, but the sequence-structure-function relations remain poorly characterized. In this work, we employ coarse-grained molecular dynamics simulations within an active learning protocol employing deep representational learning and Bayesian optimization to efficiently identify molecules capable of assembling pseudo-1D nanoaggregates with good stacking of the electronically-active $\pi$-cores. We consider the DXXX-OPV3-XXXD oligopeptide family, where D is an Asp residue and OPV3 is an oligophenylene vinylene oligomer (1,4-distyrylbenzene), to identify the top performing XXX tripeptides within all 20$^3$ = 8,000 possible sequences. By direct simulation of only 2.3% of this space, we identify molecules predicted to exhibit superior assembly relative to those reported in prior work. Spectral clustering of the top candidates reveals new design rules governing assembly. This work establishes new understanding of DXXX-OPV3-XXXD assembly, identifies promising new candidates for experimental testing, and presents a computational design platform that can be generically extended to other peptide-based and peptide-like systems.

研究动机与目标

  • 高效探索包含8,000个成员的DXXX-OPV3-XXXD序列空间,以寻找最优自组装π-共轭肽。
  • 通过计算模拟优先筛选最有希望的候选者,以克服实验筛选的高昂成本。
  • 识别控制纳米聚集体组装质量与堆叠有序度的序列-结构-功能关系。
  • 开发一种可推广的计算平台,用于设计基于肽及类肽的功能材料。

提出的方法

  • 采用粗粒度分子动力学(CGMD)模拟,在微秒时间尺度上模拟DXXX-OPV3-XXXD肽的自组装过程。
  • 利用变分自编码器(VAEs)学习高维构象轨迹的低维表征。
  • 应用高斯过程回归,基于模拟数据构建组装质量的代理模型。
  • 利用贝叶斯优化迭代选择最具有信息量的序列进行下一轮模拟,以最小化总模拟成本。
  • 集成主动学习循环,持续优化预测结果,并在最小采样量下收敛至性能最优的候选序列。
  • 对轨迹执行谱聚类,以识别不同的组装路径,并根据结构与动力学特征将序列分类为优秀、中等或较差的自组装体。

实验结果

研究问题

  • RQ1哪些DXXX-OPV3-XXXD三肽序列能形成最有序、高度堆叠的纳米聚集体,并实现最优π轨道重叠?
  • RQ2如何高效导航包含8,000个成员的序列空间,在不进行 exhaustive 模拟的前提下识别出最优候选者?
  • RQ3哪些物理化学特性或序列模式与高组装质量及结构有序度相关?
  • RQ4调控π-共轭肽自组装为功能性光电器件纳米结构的设计原理是什么?
  • RQ5结合CGMD、表征学习与贝叶斯优化的主动学习框架,能否以极低模拟成本可靠预测性能最优的候选者?

主要发现

  • 主动学习协议在仅模拟8,000种可能三肽序列中的186个(占全空间的2.3%)后,成功识别出性能最优的DXXX-OPV3-XXXD序列。
  • 该方法预测的组装质量优于先前报道的实验候选者,表明其在堆叠与纳米聚集体形成方面性能更优。
  • 对模拟轨迹执行谱聚类揭示了低维流形结构,并根据结构与动力学特征自然地将序列划分为优秀、中等与较差的自组装体。
  • 优秀自组装体富含小至中等大小的疏水性残基(如Ala、Val、Leu、Ile),而大体积芳香族残基(如Phe、Tyr、Trp)则显著匮乏。
  • 位于靠近π-核心的X位置的甲硫氨酸被发现对自组装具有中等至较强的促进作用,而其他位置的Asp、Glu和Met则影响甚微。
  • 本研究建立了一个计算高效、可扩展的虚拟筛选平台,适用于肽基材料,且可推广至其他π-共轭体系(如PDI-或OT-基肽)。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。