[논문 리뷰] DeepSymmetry : Using 3D convolutional networks for identification of tandem repeats and internal symmetries in protein structures
DeepSymmetry는 단일 반복 구조와 단백질 구조 내부의 대칭성을 탐지하기 위해 6차원 Veronese 임베딩을 학습하여 대칭 축을 1도 미만의 중앙각 오차로 예측하는 3D 컨volution 신경망을 소개한다. 이 방법은 높은 정확도로 대칭 순서와 축을 식별하며, RepeatsDB에 등재되지 않은 10,000개 이상의 잠재적 단일 반복 단백질을 발견한다.
Motivation: Thanks to the recent advances in structural biology, nowadays three-dimensional structures of various proteins are solved on a routine basis. A large portion of these contain structural repetitions or internal symmetries. To understand the evolution mechanisms of these proteins and how structural repetitions affect the protein function, we need to be able to detect such proteins very robustly. As deep learning is particularly suited to deal with spatially organized data, we applied it to the detection of proteins with structural repetitions. Results: We present DeepSymmetry, a versatile method based on three-dimensional (3D) convolutional networks that detects structural repetitions in proteins and their density maps. Our method is designed to identify tandem repeat proteins, proteins with internal symmetries, symmetries in the raw density maps, their symmetry order, and also the corresponding symmetry axes. Detection of symmetry axes is based on learning six-dimensional Veronese mappings of 3D vectors, and the median angular error of axis determination is less than one degree. We demonstrate the capabilities of our method on benchmarks with tandem repeated proteins and also with symmetrical assemblies. For example, we have discovered over 10,000 putative tandem repeat proteins that are not currently present in the RepeatsDB database. Availability: The method is available at https://team.inria.fr/nano-d/software/deepsymmetry. It consists of a C++ executable that transforms molecular structures into volumetric density maps, and a Python code based on the TensorFlow framework for applying the DeepSymmetry model to these maps.
연구 동기 및 목표
- 단백질 3D 구조 및 밀도도에서 구조적 반복과 내부 대칭성을 탐지하기 위한 강력한 방법을 개발하는 것.
- 3D 공간에서 연속적이고 유일한 표현이 필요한 3D 대칭 축을 예측하는 과제를 해결하는 것.
- 단백질의 부피 표현을 이용하여 단일 반복 단백질과 대칭 단백질 조립체를 식별하는 것.
- 저해상도 냉동전자현미경 밀도도에서 대칭성을 탐지할 수 있도록 하는 것.
- 기존 데이터베이스(예: RepeatsDB)를 초월하여 이전에 발견되지 않은 단일 반복 단백질을 발견하는 것.
제안 방법
- 이 방법은 PDB 구조와 냉동전자현미경 맵에서 유도된 부피 밀도도를 처리하기 위해 3D 컨volution 신경망(3D-CNN)을 사용한다.
- 3D 벡터를 연속적이고 유일하게 대칭 축 예측이 가능한 방식으로 표현하기 위해 6차원 Veronese 임베딩을 활용한다.
- 네트워크는 대칭 순서와 대칭 축 방향을 모두 예측하며, 후자는 Veronese 매핑을 사용한 각도 회귀를 통해 학습된다.
- 입력 데이터는 PDB 파일을 24×24×24 밴델스 해상도로 3D 밀도도로 변환하는 C++ 파이프라인을 통해 생성된다.
- 모델는 알려진 대칭성을 가진 합성 데이터를 사용하여 훈련되며, 성능은 별도의 검증 세트를 통해 검증된다.
- 딥러닝을 위한 TensorFlow와 데이터 전처리를 위한 C++를 사용하여 대규모 단백질 구조의 효율적 처리를 가능하게 한다.
실험 결과
연구 질문
- RQ13D 컨볼루션 신경망은 단백질 3D 구조에서 내부 대칭성과 단일 반복을 효과적으로 탐지할 수 있는가?
- RQ2비유일성과 비연속적 표현의 과제가 존재하는 3D 공간에서 대칭 축 방향을 얼마나 정확하게 예측할 수 있는가?
- RQ3저해상도 냉동전자현미경 밀도도에서 신호 대 잡음 비율이 낮은 상황에서도 대칭성을 탐지할 수 있는가?
- RQ4이 방법은 기존 데이터베이스(예: RepeatsDB)에 존재하지 않는 새로운 단일 반복 단백질을 어느 정도 발견할 수 있는가?
- RQ5모델는 다양한 대칭 순서와 구조적 맥락에서 일반화되는가?
주요 결과
- 대칭 축 예측의 중앙각 오차는 1도 미만으로, 축 탐지의 높은 정밀도를 나타낸다.
- 검증 세트에서 합성 데이터를 사용하여 대칭 순서를 정확히 예측한 정확도는 92%에 이른다.
- 검증 세트에서 모델는 평균 각도 차이 약 8°로 대칭 축을 탐지하며, 잘못된 예측 사례도 포함되어 있다.
- RepeatsDB 데이터베이스에 현재 등재되어 있지 않은 10,000개 이상의 잠재적 단일 반복 단백질이 발견되었다.
- 10개의 테스트 케이스를 통해 저해상도 냉동전자현미경 밀도도에서 대칭성을 성공적으로 식별하였다.
- 첫 번째 레이어의 3D 컨볼루션 커널은 라플라시안과 같은 기본 필터를 띠며, 두 번째 레이어의 커널은 구면 조화함수와 같은 수직 기저 집합을 띠어 계층적 특징 학습을 나타낸다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.