[论文解读] The relation between alignment covariance and background-averaged epistasis
本文建立了基于对齐的协方差与背景平均表型互作之间的数学联系,表明多序列对齐中的协方差意味着表型互作效应,但反之不成立。通过重新参数化序列数据并应用逆沃尔什-哈达玛变换到对齐统计量,作者证明了即使在数据深度有限的情况下,也能从天然对齐中准确重建组合突变体的功能预测,与实验数据的斯皮尔曼等级相关系数高达 R² = 0.86。
Epistasis, or the context-dependence of the effects of mutations, limits our ability to predict the functional impact of combinations of mutations, and ultimately our ability to predict evolutionary trajectories. Information about the context-dependence of mutations can essentially be obtained in two ways: First, by experimental measurement the functional effects of combinations of mutations and calculating the epistatic contributions directly, and second, by statistical analysis of the frequencies and co-occurrences of protein residues in a multiple sequence alignment of protein homologs. In this manuscript, we derive the mathematical relationship between epistasis calculated on the basis of functional measurements, and the covariance calculated from a multiple sequence alignment. There is no one-to-one mapping between covariance and epistatic terms: covariance implies epistasis, but epistasis does not necessarily lead to covariance, indicating that covariance in itself is not the directly relevant quantity for functional prediction. Having calculated epistatic contributions from the alignment, we can directly obtain a functional prediction from the alignment statistics by applying a Walsh-Hadamard transform, fully analogous to the transformation that reconstructs functional data from measured epistatic contributions. This embedding into the Hadamard framework is directly relevant for solidifying our theoretical understanding of statistical methods that predict function and three-dimensional structure from natural alignments.
研究动机与目标
- 建立多序列对齐中统计模式与功能表型互作之间理论框架的理论基础。
- 弥合基于协方差推断与蛋白质序列中实验测量表型互作之间的概念鸿沟。
- 仅使用对齐统计量,无需实验检测,实现对组合突变体的功能预测。
- 验证基于对齐的表型互作项可通过逆哈达玛变换高精度重构表型。
提出的方法
- 假设二元表型(功能/非功能),将对齐频率映射到表型互作贡献。
- 将基因型从 {0,1} 重新参数化为 {-1,1},以匹配背景平均表型互作形式。
- 通过将对齐列的逐元素乘积的均值乘以序列空间大小,定义基于对齐的表型互作项。
- 应用逆沃尔什-哈达玛变换,从基于对齐的表型互作项重构表型值。
- 使用逆变换 $\boldsymbol{\hat{y}}^{\mathrm{aln}} = \boldsymbol{H}^{-1}\boldsymbol{V}^{-1}\boldsymbol{\bar{\omega}}^{\mathrm{aln}}$ 预测功能结果。
- 将基于对齐统计量的预测结果与荧光蛋白 2^13 突变体文库的实验表型互作数据进行比较。
实验结果
研究问题
- RQ1在蛋白质序列中,对齐协方差与背景平均表型互作之间在数学上如何关联?
- RQ2是否可仅从天然序列对齐中准确推导出组合突变体的功能预测?
- RQ3协方差与表型互作之间不存在一一对应关系,是否限制了基于对齐方法的预测能力?
- RQ4在对齐深度有限的情况下,其表型互作估计在多大程度上仍能用于可靠的功能预测?
- RQ5哈达玛框架能否扩展到具有 20 种氨基酸选择的天然蛋白质序列空间?
主要发现
- 多序列对齐中的协方差意味着表型互作,但表型互作不一定会导致可检测到的协方差,表明协方差并非功能上下文依赖性的直接代理。
- 基于对齐的表型互作项(通过重参数化序列的列均值计算)可通过逆沃尔什-哈达玛变换用于重构表型值。
- 即使仅使用总共 8,192 种可能突变体中的 200–300 个序列,该方法在预测与测量表型之间的斯皮尔曼等级相关系数 ρ ≈ 0.9 时仍表现优异。
- 对于完整的功能突变体数据集,基于对齐的表型互作与实验测量表型互作的 R² 值为 0.86(一阶与二阶项合并)。
- 当对齐深度受限时,该方法依然稳健,表明功能序列的代表性子集足以实现准确预测。
- 该理论框架为利用天然序列对齐预测功能与三维结构提供了坚实基础,尽管扩展至 20 种氨基酸状态仍是开放挑战。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。