Skip to main content
QUICK REVIEW

[论文解读] p-Adic numbers in bioinformatics: from genetic code to PAM-matrix

Andrei Yu. Khrennikov, С. В. Козырев|arXiv (Cornell University)|Mar 1, 2009
advanced mathematical theories参考文献 16被引用 3
一句话总结

本文提出了一种新颖的2-adic数论框架,用于分析生物信息学中的遗传密码与PAM矩阵。通过使用2-adic距离重新参数化氨基酸索引,PAM250矩阵被分解为一个2-adically正则分量和一个稀疏偏差矩阵,结果表明非零偏差与芳香族氨基酸(Y, W, F)及半胱氨酸(C)强烈相关,暗示其结构与化学性质破坏了2-adic正则性。

ABSTRACT

In this paper we denonstrate that the use of the system of 2-adic numbers provides a new insight to some problems of genetics, in particular, generacy of the genetic code and the structure of the PAM matrix in bioinformatics. The 2-adic distance is an ultrametric and applications of ultrametrics in bioinformatics are not surprising. However, by using the 2-adic numbers we match ultrametric with a number theoretic structure. In this way we find new applications of an ultrametric which differ from known up to now in bioinformatics. We obtain the following results. We show that the PAM matrix A allows the expansion into the sum of the two matrices A=A^{(2)}+A^{(\infty)}, where the matrix A^{(2)} is 2-adically regular (i.e. matrix elements of this matrix are close to locally constant with respect to the discussed earlier by the authors 2-adic parametrization of the genetic code), and the matrix A^{(\infty)} is sparse. We discuss the structure of the matrix A^{(\infty)} in relation to the side chain properties of the corresponding amino acids.

研究动机与目标

  • 探究p-adic数论,特别是2-adic结构,是否能为遗传密码与PAM矩阵的组织方式提供新见解。
  • 检验PAM矩阵结构是否反映超出标准超度量聚类的层级化与数论模式。
  • 将PAM矩阵分解为分离2-adically正则行为与与特定氨基酸性质相关的局部偏差的分量。
  • 将观察到的偏离2-adic正则性的偏差与氨基酸的化学与几何特征(尤其是侧链性质)联系起来。

提出的方法

  • 利用遗传密码的简并性,通过2-adic二维映射重新参数化20种氨基酸,实现分层的、局部恒定的表示。
  • 将此2-adic参数化应用于PAM250矩阵的行与列重索引,以揭示取代概率中潜在的2-adic正则性。
  • 将PAM矩阵分解为 A = A^(2) + A^(∞),其中 A^(2) 为2-adically正则(局部恒定),A^(∞) 为稀疏矩阵。
  • 分析 A^(∞) 中非零元素的分布,识别对偏离2-adic正则性贡献最大的氨基酸。
  • 将 A^(∞) 中非零元素的位置与氨基酸的已知生化性质(如芳香性与二硫键形成能力)进行相关性分析。
  • 利用2-adic度量的强三角不等式,解释氨基酸在蛋白质进化与序列比对背景下的分层聚类。

实验结果

研究问题

  • RQ1能否通过2-adic数论有意义地重构PAM矩阵,以揭示氨基酸取代模式中的隐藏正则性?
  • RQ22-adic参数化对遗传密码的解释程度在多大程度上能将PAM矩阵结构解释为局部恒定函数?
  • RQ3哪些氨基酸主要导致PAM矩阵中偏离2-adic正则性?其生化性质如何与这些偏差相关?
  • RQ4A^(∞)矩阵的稀疏性与特定氨基酸(如酪氨酸、色氨酸、苯丙氨酸与半胱氨酸)的侧链特性之间有何关联?
  • RQ52-adic框架能否为PAM矩阵中编码的进化动力学提供新的理论基础?

主要发现

  • PAM250矩阵可分解为 A = A^(2) + A^(∞),其中 A^(2) 为2-adically正则(局部恒定),A^(∞) 为稀疏矩阵,表明矩阵的大部分结构由2-adic正则性捕获。
  • A^(∞) 中的非零元素主要位于对应于酪氨酸(Y)、色氨酸(W)、苯丙氨酸(F)和半胱氨酸(C)的索引处,此外精氨酸(R)也有额外贡献。
  • A^(∞) 中偏离2-adic正则性的偏差与芳香族侧链(Y, W, F)的存在以及半胱氨酸(C)形成二硫键的能力强烈相关。
  • 疏水性氨基酸聚集在两个2-adic球内,表明2-adic结构自然地按疏水性与化学类别对氨基酸进行分组。
  • 矩阵元素 A^(∞)_{FY} 尤其大且为正,表明苯丙氨酸与酪氨酸之间存在强烈的偏离2-adic正则性,二者为结构相似的芳香族氨基酸。
  • 2-adic参数化为生物信息学中的超度量结构提供了数论基础,将p-adic分析与蛋白质进化及序列比对联系起来。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。