[论文解读] Exploring the consequences of lack of closure in codon models
本研究探讨了在进化过程异质性条件下,密码子模型缺乏数学封闭性如何导致对选择压力($\omega$)和分支长度的估计偏差。通过模拟实验表明,即使底层DNA模型是封闭的,广泛用于系统发育分析的非封闭密码子模型仍可能使$\omega$的估计误差高达17%,分支长度误差超过50%,这表明遗传密码的扩展引入了非线性,从而破坏了模型的准确性。
Models of codon evolution are commonly used to identify positive selection. Positive selection is typically a heterogeneous process, i.e., it acts on some branches of the evolutionary tree and not others. Previous work on DNA models showed that when evolution occurs under a heterogeneous process it is important to consider the property of model closure, because non-closed models can give biased estimates of evolutionary processes. The existing codon models that account for the genetic code are not closed; to establish this it is enough to show that they are not linear (meaning that the sum of two codon rate matrices in the model is not a matrix in the model). This raises the concern that a single codon model fit to a heterogeneous process might mis-estimate both the effect of selection and branch lengths. Codon models are typically constructed by choosing an underlying DNA model (e.g., HKY) that acts identically and independently at each codon position, and then applying the genetic code via the parameter $ω$ to modify the rate of transitions between codons that code for different amino acids. Here we use simulation to investigate the accuracy of estimation of both the selection parameter $ω$ and branch lengths in cases where the underlying DNA process is heterogeneous but $ω$ is constant. We find that both $ω$ and branch lengths can be mis-estimated in these scenarios. Errors in $ω$ were usually less than 2% but could be as high as 17%. We also assessed if choosing different underlying DNA models had any affect on accuracy, in particular we assessed if using closed DNA models gave any advantage. However, a DNA model being closed does not imply that the codon model constructed from it is closed, and in general we found that using closed DNA models did not decrease errors in the estimation of $ω$.
研究动机与目标
- 研究模型非封闭性对异质进化过程中密码子模型参数估计准确性的影响。
- 确定使用封闭DNA模型(例如Lie-Markov模型)是否能提高密码子模型的估计准确性。
- 评估密码子模型中估计偏差的主要来源是否源于遗传密码扩展参数引入的非线性。
- 探索是否可以构建一个封闭且具有生物学意义的密码子模型,或此类模型仅存在过于复杂的非实用替代方案。
- 提出一种保持遗传密码结构的线性密码子模型,同时减少由非封闭性引起的偏差。
提出的方法
- 在已知进化过程下,模拟非平稳、非同质的密码子序列数据,其$\omega$和分支长度各不相同。
- 在同质、时间可逆的假设下,将标准密码子模型(MG型和GY型)拟合到模拟数据上。
- 比较不同底层DNA模型(包括封闭的Lie-Markov模型和非封闭的GTR、F81等模型)下$\omega$和分支长度的估计误差。
- 通过检查模型中速率矩阵之和是否仍属于该模型(即线性性)来分析密码子模型的数学结构,这是实现封闭性的先决条件。
- 通过将非线性约束($\alpha_i, \alpha_i\omega$)替换为独立参数($\alpha_i, \mu_i$),构建包含MG型模型的最小线性密码子模型,从而实现$\omega_i = \mu_i / \alpha_i$。
- 在模拟中评估所构建线性模型的性能,以评估其是否能降低估计偏差。
实验结果
研究问题
- RQ1在异质进化过程中,密码子模型的非封闭性在多大程度上导致对选择参数$\omega$的估计偏差?
- RQ2使用封闭DNA模型(如Lie-Markov模型)是否能提高密码子模型中$\omega$和分支长度估计的准确性?
- RQ3密码子模型中估计偏差的主要来源是否可归因于遗传密码扩展参数引入的非线性?
- RQ4能否构建一个具有生物学意义且封闭的密码子模型,还是此类模型过于复杂而不具实用性?
- RQ5将密码子模型中的非线性约束替换为独立参数,是否能降低估计偏差,同时保持生物学可解释性?
主要发现
- 即使底层DNA模型是封闭的,由于遗传密码引入的非线性,非封闭密码子模型对$\omega$的估计误差仍可能高达17%。
- 在异质进化过程中,非封闭密码子模型对分支长度的估计误差可超过50%。
- 使用封闭DNA模型(如Lie-Markov模型)并不能降低密码子模型中的估计误差,表明偏差的根源在于密码子层面的扩展,而非DNA模型的封闭性。
- 即使最通用的非平稳密码子模型(GNC)无法通过绝对拟合优度检验,其原因可能在于密码子模型缺乏封闭性。
- 在不包含$G$矩阵(即无遗传密码扩展)的模拟中,误差小了一个数量级,证实遗传密码的非线性约束是主要偏差来源。
- 可以构建一个保持遗传密码结构且允许$\omega_i = \mu_i / \alpha_i$的最小线性密码子模型,尽管其需要八个参数而非五个,但可能为减少偏差提供可行路径。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。