[Paper Review] Protein Structure and Sequence Generation with Equivariant Denoising Diffusion Probabilistic Models
A fully data-driven diffusion model generating large protein structures, sequences, and rotamers that are physically plausible and conditionable on compact topology constraints.
Proteins are macromolecules that mediate a significant fraction of the cellular processes that underlie life. An important task in bioengineering is designing proteins with specific 3D structures and chemical properties which enable targeted functions. To this end, we introduce a generative model of both protein structure and sequence that can operate at significantly larger scales than previous molecular generative modeling approaches. The model is learned entirely from experimental data and conditions its generation on a compact specification of protein topology to produce a full-atom backbone configuration as well as sequence and side-chain predictions. We demonstrate the quality of the model via qualitative and quantitative analysis of its samples. Videos of sampling trajectories are available at https://nanand2.github.io/proteins .
Motivation & Objective
- Motivate the design of proteins with specific 3D structures and chemical properties at scale.
- Present a fully data-driven generative model for protein structure, sequence, and rotamers.
- Enable generation conditioned on compact topology constraints to produce diverse, physically plausible proteins.
Proposed method
- Use a denoising diffusion probabilistic model trained on experimental protein data to generate backbone coordinates, rotations, sequence, and side-chain torsions.
- Employ equivariant Transformers with invariant point attention to ensure rotational/ translational invariance and equivariance.
- Diffuse rotations by interpolating on SU(2) using SLERP and diffuse discrete sequences via a masked language-model–like discrete diffusion.
- Train separate diffusion models for structure, sequence, and rotamers, with a compact constraint conditioning scheme that encodes topology at the block level.
- Incorporate frame-aligned point error (FAPE) loss to train the rotationally invariant denoiser.
Experimental results
Research questions
- RQ1Can a diffusion model generate large, physically plausible protein structures across diverse PDB domain topologies?
- RQ2How well can the model design sequences and pack rotamers conditioned on generated backbones?
- RQ3To what extent can compact topology constraints guide controllable protein generation and inpainting?
- RQ4What is the impact of joint versus separate diffusion modeling for structure and sequence?
Key findings
- The model produces high-quality, diverse protein structures with realistic hydrogen bond patterns and backbone geometry.
- Sequence design and rotamer packing are comparable to or faster than baselines and show competitive recovery rates.
- The model supports inpainting and controllable design, including topology modification, loop design, and variable-length loops.
- Joint modeling of structure and sequence demonstrates contextually plausible inpainting and potential for full-atom loop and Ig design.
- Samples closely reflect biophysical priors such as Ramachandran distributions and bond length/angle histograms.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.