Skip to main content
QUICK REVIEW

[논문 리뷰] Persistent cohomology for data with multicomponent heterogeneous information

Zixuan Cang, Guo‐Wei Wei|arXiv (Cornell University)|2018. 07. 29.
Topological and Geometric Data Analysis참고 문헌 41인용 수 7
한 줄 요약

이 논문은 기하적 구조와 함께 원자 전하 및 전기장 잠재 에너지와 같은 다성분 이질적 자료를 위상적 불변량에 체계적으로 통합하는 영구 코homology 프레임워크를 제안한다. 단체 복합체 위에서 스무스한 1-코사이클을 계산함으로써, 표준 영구 바코드에 물리적 정보를 추가함으로써 단백질-리간드 결합 친화도 예측 정확도를 크게 향상시킨다. 특히 전기적 성질을 통합할 경우 성능 향상이 두드러진다.

ABSTRACT

Persistent homology is a powerful tool for characterizing the topology of a data set at various geometric scales. When applied to the description of molecular structures, persistent homology can capture the multiscale geometric features and reveal certain interaction patterns in terms of topological invariants. However, in addition to the geometric information, there is a wide variety of non-geometric information of molecular structures, such as element types, atomic partial charges, atomic pairwise interactions, and electrostatic potential function, that is not described by persistent homology. Although element specific homology and electrostatic persistent homology can encode some non-geometric information into geometry based topological invariants, it is desirable to have a mathematical framework to systematically embed both geometric and non-geometric information, i.e., multicomponent heterogeneous information, into unified topological descriptions. To this end, we propose a mathematical framework based on persistent cohomology. In our framework, non-geometric information can be either distributed globally or resided locally on the datasets in the geometric sense and can be properly defined on topological spaces, i.e., simplicial complexes. Using the proposed persistent cohomology based framework, enriched barcodes are extracted from datasets to represent heterogeneous information. We consider a variety of datasets to validate the present formulation and illustrate the usefulness of the proposed persistent cohomology. It is found that the proposed framework using cohomology boosts the performance of persistent homology based methods in the protein-ligand binding affinity prediction on massive biomolecular datasets.

연구 동기 및 목표

  • 영구 호모로지가 원자 전하 및 전기장 잠재 에너지와 같은 비기하학적 분자 성질을 포착하지 못하는 한계를 해결하기 위해.
  • 기하학적 및 비기하학적 정보를 위상적 데이터 분석에서 통합하는 수학적 프레임워크를 개발하기 위해.
  • 원소 종류, 부분 전하, 상호작용 에너지와 같은 이질적 자료를 위상 기반 기술자에 체계적으로 통합할 수 있도록 하기 위해.
  • 특히 단백질-리간드 결합 친화도 예측에 있어 위상적 방법의 예측 능력을 향상시키기 위해.

제안 방법

  • 거리 또는 척도 파rameter에 기반한 필터레이션을 사용하여 점 클러스터 데이터에서 단체 복합체를 구축하기.
  • 단체 복합체 위에 가중치가 부여된 그래프 라플라시안을 정의하여 비기하학적 자료를 나타내는 스무스한 1-코사이클을 계산하기.
  • 스무스한 1-코사이클을 단체의 함수로 사용하여 표준 영구 바코드에 물리적 정보를 추가하기.
  • 기하학적 및 비기하학적 자료를 모두 포함하는 enriched barcodes를 비교하기 위해 수정된 워싱어스타인 거리 도입하기.
  • 정점에 전기장 잠재 에너지 값을 할당하고 코homology를 통해 이를 전파함으로써 생분자 데이터셋에 프레임워크 적용하기.
  • 그리드 서치를 통해 초모수를 최적화한 기울기 부스팅을 사용하여 enriched barcodes에서 결합 친화도 예측하기.

실험 결과

연구 질문

  • RQ1영구 코hom로지가 원자 부분 전하와 같은 비기하학적 분자 성질을 위상적 불변량에 통합하는 데 사용될 수 있는가?
  • RQ2물리적 자료로 바코드를 강화할 경우 기계학습 모델의 단백질-리간드 결합 친화도 예측 성능에 어떤 영향을 미치는가?
  • RQ3코homology 기반 기술자로는 동질성보다도 고리나 빈 공간과 같은 위상적 특징에 물리적 성질을 더 잘 국소화하고 연관시킬 수 있는가?
  • RQ4대규모 생분자 데이터셋에서 전기적 정보를 통합할 경우 위상 모델의 예측 정확도가 어느 정도 향상되는가?
  • RQ5제안된 프레임워크는 반데르발스 상호작용이나 원자 상호작용 에너지와 같은 다른 물리적 성질로 일반화 가능한가?

주요 결과

  • 영구 코hom로지 프레임워크는 단체 복합체 위에서 스무스한 1-코사이클을 통해 전기장 잠재 에너지와 같은 비기하학적 정보를 위상적 불변량에 성공적으로 통합하였다.
  • 바코드에 전기적 정보를 통합함으로써, 테스트한 PDBbind 모든 버전에서 단백질-리간드 결합 친화도 예측 성능이 향상되었으며, 특히 v2016에서 가장 높은 향상이 관찰되었다 (피어슨 상관계수: 0.778 vs. 0.767, 표준 영구 호모로지 대비).
  • 영구 코호모로지에 전기적 성질을 통합함으로써, PDBbind v2016 코어 세트에서 중앙값 피어슨 상관계수 0.833 (pKd)를 기록하여 거의 최적의 성능을 달성하였다.
  • 수정된 워싱어스타인 거리는 enriched barcodes 간의 유사성을 효과적으로 측정하여, 이질적 자료를 포함한 위상 기술자의 견고한 비교를 가능하게 하였다.
  • 0차 및 고차원 영구 코호모로지 모두에서 표준 영구 호모로지보다 일관된 성능 향상을 보였으며, 특히 물리적 상호작용과 관련된 생물학적으로 의미 있는 특징을 잘 포착하였다.
  • 이 방법은 전기적 성질을 넘어서 반데르발스 상호작용 등 다른 물리적 성질로도 일반화 가능하며, 더 풍부한 위상 모델링을 위한 복잡한 데이터셋의 표현에 기여한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.