Skip to main content
QUICK REVIEW

[논문 리뷰] Inferring a Continuous Distribution of Atom Coordinates from Cryo-EM Images using VAEs

Dan Rosenbaum|arXiv (Cornell University)|2021. 06. 26.
Advanced Electron Microscopy Techniques and Applications참고 문헌 24인용 수 31
한 줄 요약

논문은 conformation과 pose를 분리된 잠재 공간에 인코딩하고 차등 가능한 cryo-EM 렌더러를 통해 원자 좌표를 렌더링함으로써 cryo-EM 입자 이미지로부터 연속적인 원자 단백질 구조 분포를 직접 재구성하는 완전한 엔드-투-엔드 변분 자동인코더 프레임워크를 소개한다.

ABSTRACT

Cryo-electron microscopy (cryo-EM) has revolutionized experimental protein structure determination. Despite advances in high resolution reconstruction, a majority of cryo-EM experiments provide either a single state of the studied macromolecule, or a relatively small number of its conformations. This reduces the effectiveness of the technique for proteins with flexible regions, which are known to play a key role in protein function. Recent methods for capturing conformational heterogeneity in cryo-EM data model it in volume space, making recovery of continuous atomic structures challenging. Here we present a fully deep-learning-based approach using variational auto-encoders (VAEs) to recover a continuous distribution of atomic protein structures and poses directly from picked particle images and demonstrate its efficacy on realistic simulated data. We hope that methods built on this work will allow incorporation of stronger prior information about protein structure and enable better understanding of non-rigid protein structures.

연구 동기 및 목표

  • cryo-EM 데이터에서 구성 이질성을 이산적 상태를 넘어 회복하는 것을 동기화한다.
  • 노이즈가 있는 2D 투사로부터 원자 구조의 분포를 추론하는 엔드-투-엔드 딥러닝 모델을 제안한다.
  • 학습을 위해 원자 좌표를 시뮬레이션된 EM 이미지로 매핑하는 차등 가능한 렌더러를 활용한다.
  • 재구성을 규제하기 위해 결합 길이/ 기하학과 같은 명시적 원자 공간 선험치를 통합한다.
  • 다중 모달 회복과 연속 잠재 표현을 입증하기 위해 시뮬레이션 데이터로 평가한다.

제안 방법

  • 원자 좌표에서 EM 이미지로의 이미징 형성을 차등 가능한 렌더러로 모델링한다.
  • 구조와 방향을 분리하기 위해 conformation(32-d)과 pose(32-d)의 두 잠재공간을 가진 변분 자동인코더를 사용한다.
  • 잠재 샘플을 기본 구성 대비 delta 백본 프레임(잔류별 자세)으로 해독한다.
  • 해독된 원자 구성들을 평균 이미지로 렌더링하고 관측 잡음을 모델링한다.
  • 이미지당 여러 posterior z를 샘플링하여 다중 모드 커버리지를 촉진하는 IWAE와 유사한 목표로 학습한다.
  • 백본 연속성 손실과 중심화 손실 등 보조 손실을 포함하고; 회전 행렬을 얻기 위해 Gram-Schmidt를 사용한다; 배치 단위로 Adam으로 최적화한다.

실험 결과

연구 질문

  • RQ1Can a VAE-based model recover a continuous distribution of atomic protein conformations directly from 2D cryo-EM images?
  • RQ2Does disentangling pose and conformation in latent space improve recovery of multimodal structural states compared to a single latent space?
  • RQ3Can a differentiable renderer and atom-space priors enable end-to-end learning from images to atom coordinates?
  • RQ4How well does the method capture multimodal conformational distributions and what are its limitations on simulated data?

주요 결과

  • The method can train neural networks to represent diverse states of a protein from 2D particle images (AurA kinase) with a true distribution approximated within 3.35 Angstrom by the EMD-RMSD metric.
  • Latent space interpolation shows continuous transitions between states, demonstrating a continuous distribution of conformations.
  • Disentangling pose and conformation latent spaces and pre-training pose prediction are crucial to avoid degenerate solutions and enable non-degenerate structural reconstructions.
  • Conditional sampling from the posterior yields weaker correlations to the true distributions under high-noise 2D data, indicating the necessity of shared capacity across many images.
  • The evaluation on simulated data acknowledges limitations, with the model performing well in unconditional sampling and partial success in conditional predictions.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.