[Paper Review] p-Adic numbers in bioinformatics: from genetic code to PAM-matrix
This paper introduces a novel 2-adic number-theoretic framework to analyze the genetic code and PAM matrix in bioinformatics. By reparametrizing amino acid indices using 2-adic distance, the PAM250 matrix is decomposed into a 2-adically regular component and a sparse deviation matrix, revealing that non-zero deviations correlate strongly with aromatic amino acids (Y, W, F) and cysteine (C), suggesting their structural and chemical properties disrupt 2-adic regularity.
In this paper we denonstrate that the use of the system of 2-adic numbers provides a new insight to some problems of genetics, in particular, generacy of the genetic code and the structure of the PAM matrix in bioinformatics. The 2-adic distance is an ultrametric and applications of ultrametrics in bioinformatics are not surprising. However, by using the 2-adic numbers we match ultrametric with a number theoretic structure. In this way we find new applications of an ultrametric which differ from known up to now in bioinformatics. We obtain the following results. We show that the PAM matrix A allows the expansion into the sum of the two matrices A=A^{(2)}+A^{(\infty)}, where the matrix A^{(2)} is 2-adically regular (i.e. matrix elements of this matrix are close to locally constant with respect to the discussed earlier by the authors 2-adic parametrization of the genetic code), and the matrix A^{(\infty)} is sparse. We discuss the structure of the matrix A^{(\infty)} in relation to the side chain properties of the corresponding amino acids.
Motivation & Objective
- To investigate whether p-adic number theory, particularly 2-adic structures, can provide new insights into the organization of the genetic code and PAM matrix.
- To test the hypothesis that the PAM matrix’s structure reflects underlying hierarchical and number-theoretic patterns beyond standard ultrametric clustering.
- To decompose the PAM matrix into components that separate 2-adically regular behavior from localized deviations tied to specific amino acid properties.
- To link the observed deviations from 2-adic regularity to the chemical and geometric features of amino acids, especially side chain properties.
Proposed method
- Reparametrize the 20 amino acids using a 2-adic 2-dimensional mapping derived from the genetic code’s degeneracy, enabling a hierarchical, locally constant representation.
- Apply this 2-adic parametrization to reindex the rows and columns of the PAM250 matrix to reveal underlying 2-adic regularity in substitution probabilities.
- Decompose the PAM matrix as A = A^(2) + A^(∞), where A^(2) is 2-adically regular (locally constant) and A^(∞) is sparse.
- Analyze the distribution of non-zero elements in A^(∞) to identify which amino acids contribute most to deviations from 2-adic regularity.
- Correlate the positions of non-zero elements in A^(∞) with known biochemical properties of amino acids, such as aromaticity and disulfide bond formation.
- Use the strong triangle inequality of the 2-adic metric to interpret the hierarchical clustering of amino acids in the context of protein evolution and sequence alignment.
Experimental results
Research questions
- RQ1Can the PAM matrix be meaningfully restructured using 2-adic number theory to reveal hidden regularities in amino acid substitution patterns?
- RQ2To what extent does the 2-adic parametrization of the genetic code explain the structure of the PAM matrix as a locally constant function?
- RQ3Which amino acids are primarily responsible for deviations from 2-adic regularity in the PAM matrix, and how do their biochemical properties relate to these deviations?
- RQ4How does the sparsity of the A^(∞) matrix correlate with the side chain characteristics of specific amino acids like tyrosine, tryptophan, phenylalanine, and cysteine?
- RQ5Can the 2-adic framework provide a new theoretical basis for understanding the evolutionary dynamics encoded in the PAM matrix?
Key findings
- The PAM250 matrix can be decomposed into A = A^(2) + A^(∞), where A^(2) is 2-adically regular (locally constant), and A^(∞) is sparse, indicating that most of the matrix structure is captured by 2-adic regularity.
- Non-zero elements in A^(∞) are predominantly located at indices corresponding to tyrosine (Y), tryptophan (W), phenylalanine (F), and cysteine (C), with additional contributions from arginine (R).
- The deviations from 2-adic regularity in A^(∞) are strongly correlated with the presence of aromatic side chains (Y, W, F) and the ability of cysteine (C) to form disulfide bonds.
- Hydrophobic amino acids are clustered in two 2-adic balls, suggesting that 2-adic structure naturally groups amino acids by hydrophobicity and chemical class.
- The matrix element A^(∞)_{FY} is particularly large and positive, indicating a strong deviation from 2-adic regularity between phenylalanine and tyrosine, which are structurally similar aromatic amino acids.
- The 2-adic parametrization provides a number-theoretic foundation for ultrametric structures in bioinformatics, linking p-adic analysis to protein evolution and sequence alignment.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.