Skip to main content
QUICK REVIEW

[Paper Review] Discovery of Self-Assembling $\pi$-Conjugated Peptides by Active Learning-Directed Coarse-Grained Molecular Simulation

Kirill Shmilovich, Rachael A. Mansbach|arXiv (Cornell University)|Jan 27, 2020
Supramolecular Self-Assembly in Materials92 references4 citations
TL;DR

This study develops an active learning framework integrating coarse-grained molecular dynamics, variational autoencoders, and Bayesian optimization to efficiently identify top-performing π-conjugated peptides in the DXXX-OPV3-XXXD family that self-assemble into well-stacked, pseudo-1D nanoaggregates. By simulating only 2.3% of the 8,000 possible tripeptide sequences, the method identifies superior assemblers and reveals design rules favoring small-to-moderate hydrophobic residues and specific methionine positioning near the π-core.

ABSTRACT

Electronically-active organic molecules have demonstrated great promise as novel soft materials for energy harvesting and transport. Self-assembled nanoaggregates formed from $\pi$-conjugated oligopeptides composed of an aromatic core flanked by oligopeptide wings offer emergent optoelectronic properties within a water soluble and biocompatible substrate. Nanoaggregate properties can be controlled by tuning core chemistry and peptide composition, but the sequence-structure-function relations remain poorly characterized. In this work, we employ coarse-grained molecular dynamics simulations within an active learning protocol employing deep representational learning and Bayesian optimization to efficiently identify molecules capable of assembling pseudo-1D nanoaggregates with good stacking of the electronically-active $\pi$-cores. We consider the DXXX-OPV3-XXXD oligopeptide family, where D is an Asp residue and OPV3 is an oligophenylene vinylene oligomer (1,4-distyrylbenzene), to identify the top performing XXX tripeptides within all 20$^3$ = 8,000 possible sequences. By direct simulation of only 2.3% of this space, we identify molecules predicted to exhibit superior assembly relative to those reported in prior work. Spectral clustering of the top candidates reveals new design rules governing assembly. This work establishes new understanding of DXXX-OPV3-XXXD assembly, identifies promising new candidates for experimental testing, and presents a computational design platform that can be generically extended to other peptide-based and peptide-like systems.

Motivation & Objective

  • To efficiently explore the vast 8,000-member DXXX-OPV3-XXXD sequence space for optimal self-assembling π-conjugated peptides.
  • To overcome the prohibitive cost of experimental screening by using computational simulation to prioritize the most promising candidates.
  • To identify sequence-structure-function relationships governing nanoaggregate assembly quality and stacking order.
  • To develop a generalizable computational platform for designing peptide-based and peptide-like functional materials.

Proposed method

  • Employ coarse-grained molecular dynamics (CGMD) simulations to model self-assembly of DXXX-OPV3-XXXD peptides over microsecond timescales.
  • Use variational autoencoders (VAEs) to learn low-dimensional representations of high-dimensional conformational trajectories.
  • Apply Gaussian process regression to build surrogate models of assembly quality based on simulation data.
  • Utilize Bayesian optimization to iteratively select the most informative sequences to simulate next, minimizing total simulation cost.
  • Integrate active learning loops that refine predictions and converge on top-performing candidates with minimal sampling.
  • Perform spectral clustering on trajectories to identify distinct assembly pathways and classify sequences as good, intermediate, or poor assemblers.

Experimental results

Research questions

  • RQ1Which DXXX-OPV3-XXXD tripeptide sequences form the most well-ordered, highly stacked nanoaggregates with optimal π-orbital overlap?
  • RQ2How can we efficiently navigate the 8,000-member sequence space to identify top candidates without exhaustive simulation?
  • RQ3What physicochemical features or sequence patterns correlate with high assembly quality and structural order?
  • RQ4What design principles govern the self-assembly of π-conjugated peptides into functional optoelectronic nanostructures?
  • RQ5Can an active learning framework combining CGMD, representation learning, and Bayesian optimization reliably predict top performers with minimal simulation cost?

Key findings

  • The active learning protocol identified the top-performing DXXX-OPV3-XXXD sequences after simulating only 186 out of 8,000 possible tripeptide sequences, representing 2.3% of the full space.
  • The method predicted superior assembly quality compared to previously reported experimental candidates, indicating improved performance in stacking and nanoaggregate formation.
  • Spectral clustering of simulation trajectories revealed a low-dimensional manifold and naturally partitioned sequences into good, intermediate, and poor assemblers based on structural and dynamic features.
  • Good assemblers were enriched in small and intermediate-sized hydrophobic residues (e.g., Ala, Val, Leu, Ile) and depleted in large aromatic residues (e.g., Phe, Tyr, Trp).
  • Methionine in the X position closest to the π-core was found to moderately to strongly favor assembly, while Asp, Glu, and Met in other positions showed minimal influence.
  • The study established a computationally efficient, scalable platform for virtual screening of peptide-based materials, with potential extension to other π-conjugated systems like PDI- or OT-based peptides.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.