Skip to main content
QUICK REVIEW

[论文解读] A quantitative structure comparison with persistent similarity

Kelin Xia|arXiv (Cornell University)|Jul 12, 2017
Topological and Geometric Data Analysis参考文献 40被引用 3
一句话总结

本文提出了一种基于代数拓扑中持久同调的新型定量方法——持久相似性,用于生物分子结构比较。通过将拓扑条形码转换为持久贝蒂数函数(PBF),该方法实现了高效的多尺度比较,其比较依据为两个PBF之间相交区域与并集区域的面积比,与C44富勒烯异构体的曲率能差异具有高度相关性(r = -0.952)。

ABSTRACT

Biomolecular structure comparison not only reveals evolutionary relationships, but also sheds light on biological functional properties. However, traditional definitions of structure or sequence similarity always involve superposition or alignment and are computationally inefficient. In this paper, I propose a new method called persistent similarity, which is based on a newly-invented method in algebraic topology, known as persistent homology. Different from all previous topological methods, persistent homology is able to embed a geometric measurement into topological invariants, thus provides a bridge between geometry and topology. Further, with the proposed persistent Betti function (PBF), topological information derived from the persistent homology analysis can be uniquely represented by a series of continuous one-dimensional (1D) functions. In this way, any complicated biomolecular structure can be reduced to several simple 1D PBFs for comparison. Persistent similarity is then defined as the quotient of sizes of intersect areas and union areas between two correspondingly PBFs. If structures have no significant topological properties, a pseudo-barcode is introduced to insure a better comparison. Moreover, a multiscale biomolecular representation is introduced through the multiscale rigidity function. It naturally induces a multiscale persistent similarity. The multiscale persistent similarity enables an objective-oriented comparison. State differently, it facilitates the comparison of structures in any particular scale of interest. Finally, the proposed method is validated by four different cases. It is found that the persistent similarity can be used to describe the intrinsic similarities and differences between the structures very well.

研究动机与目标

  • 为克服基于叠加或序列比对的传统结构比较方法在计算效率上的不足以及对结构对齐的依赖性。
  • 开发一种基于拓扑的定量方法,能够捕捉全局结构特征,且不受局部几何噪声的影响。
  • 通过多尺度刚性函数引入多尺度表示,实现面向目标的比较。
  • 通过引入伪条形码机制,在拓扑特征微弱的情况下确保方法的鲁棒性。
  • 验证该方法检测多样化生物分子体系中内在结构相似性与差异性的能力。

提出的方法

  • 利用持久同调在过滤尺度上提取拓扑不变量(贝蒂数),将几何信息嵌入到拓扑特征中。
  • 引入持久贝蒂数函数(PBF)以连续的一维函数形式表示持久同调条形码,实现高效比较。
  • 将持久相似性定义为两个PBF之间相交区域面积与并集区域面积的比值,从而实现结构的定量比较。
  • 为拓扑特征可忽略的结构引入伪条形码,以维持可比性并避免歧义。
  • 采用多尺度刚性函数生成分层的生物分子表示,实现选择性尺度分析。
  • 通过在刚性函数的多个分辨率层级上应用基于PBF的相似性度量,推导出多尺度持久相似性。

实验结果

研究问题

  • RQ1持久同调能否在无需对齐或叠加的情况下,有效用于复杂生物分子结构的表示与比较?
  • RQ2持久相似性在检测生物分子结构之间内在拓扑相似性与差异性方面,其准确性如何?
  • RQ3该方法能否与物理性质(如富勒烯异构体的曲率能)实现高度相关?
  • RQ4引入伪条形码是否能提升在拓扑特征微弱的结构中相似性比较的鲁棒性?
  • RQ5多尺度持久相似性在多大程度上能够实现面向目标的、特定尺度的生物分子结构比较?

主要发现

  • 在C44富勒烯异构体中,持久相似性与总曲率能差异之间的皮尔逊相关系数(PCC)达到-0.952,优于距离和密度过滤方法。
  • 当以曲率能极端的异构体作为参考时,持久相似性与能量差异的PCC在大多数情况下超过0.80,表明对内在差异具有高度敏感性。
  • 该方法成功将复杂的生物分子结构简化为简单的1D PBF,实现了无需几何对齐的高效且可扩展的比较。
  • 多尺度持久相似性框架使研究人员能够聚焦于感兴趣的特定结构尺度,实现面向目标的分析。
  • 伪条形码机制有效实现了即使在拓扑特征极少的结构中也能进行有意义的相似性比较。
  • 在四个测试案例(核苷酸激酶、NMR构型、拉伸模拟、C44异构体)中的验证表明,该方法在捕捉内在结构关系方面具有出色的鲁棒性与通用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。