[论文解读] Survival analysis of DNA mutation motifs with penalized proportional hazards
本文提出SAMM,一种新颖的套索惩罚比例风险模型,用于推断B细胞受体体细胞高频突变中的DNA突变基序及其影响。通过使用马尔可夫链蒙特卡洛期望最大化(MCEM)算法处理未观测到的突变顺序和高维基序特征,SAMM在高维稀疏设置下,尤其在具有复杂上下文依赖性突变模式的情况下,相比现有方法实现了更简洁且更精确的模型。
Antibodies, an essential part of our immune system, develop through an intricate process to bind a wide array of pathogens. This process involves randomly mutating DNA sequences encoding these antibodies to find variants with improved binding, though mutations are not distributed uniformly across sequence sites. Immunologists observe this nonuniformity to be consistent with "mutation motifs", which are short DNA subsequences that affect how likely a given site is to experience a mutation. Quantifying the effect of motifs on mutation rates is challenging: a large number of possible motifs makes this statistical problem high dimensional, while the unobserved history of the mutation process leads to a nontrivial missing data problem. We introduce an $\ell_1$-penalized proportional hazards model to infer mutation motifs and their effects. In order to estimate model parameters, our method uses a Monte Carlo EM algorithm to marginalize over the unknown ordering of mutations. We show that our method performs better on simulated data compared to current methods and leads to more parsimonious models. The application of proportional hazards to mutation processes is, to our knowledge, novel and formalizes the current methods in a statistical framework that can be easily extended to analyze the effect of other biological features on mutation rates.
研究动机与目标
- 开发一种统计严谨的框架,用于估计DNA序列基序对B细胞受体体细胞高频突变率的影响。
- 解决由于未观测到的突变顺序和大量潜在基序所引发的高维与缺失数据问题。
- 与采用启发式或限制性假设的现有方法相比,提升模型的简洁性与可解释性。
- 正式化生存分析在突变过程建模中的应用,使其可扩展至其他生物特征。
- 提供一种可扩展、数据自适应的方法,在高维设置下选择相关基序并稳定估计。
提出的方法
- 提出一种半参数Cox比例风险模型,将突变视为失效事件,基序背景作为时变协变量。
- 应用l1-惩罚(套索)对1024种可能的5聚体进行特征选择,以促进稀疏且可解释的模型。
- 使用马尔可夫链蒙特卡洛期望最大化(MCEM)算法处理因未观测到的突变顺序而产生的缺失数据,在E步中通过吉布斯抽样近似不可计算的期望。
- 通过使用MCMC对所有可能的突变顺序进行边际化,整合似然函数,即使在历史信息未观测到的情况下,也能实现最大似然估计。
- 采用数据自适应方法选择模型的自由度,避免在高维设置下过拟合。
- 将框架扩展至可建模任意生物特征(而不仅限于序列基序),如位置或结构上下文。
实验结果
研究问题
- RQ1在存在未观测突变顺序和高维性的情况下,如何准确估计DNA基序的可变性?
- RQ2惩罚比例风险模型是否能在识别生物学上相关的突变基序方面优于现有方法?
- RQ3在体细胞高频突变分析中,突变顺序不确定性的引入如何影响模型准确性和参数估计?
- RQ4l1-惩罚在高维基序分析中能在多大程度上提升模型的简洁性与可解释性?
- RQ5该框架在多大程度上可推广至建模其他涉及隐藏历史的序列事件的生物过程?
主要发现
- SAMM仅估计了1024种可能5聚体中的137个独特基序参数,相比SHazaM(1015个)和逻辑回归(485个)展现出更优的模型简洁性。
- 在模拟数据中,SAMM在识别真实突变基序方面优于当前最先进的方法,尤其在高突变率和复杂上下文依赖性条件下。
- 该方法生成的热点与冷点模式在视觉上与其他模型相似,但结构更稀疏、更具可解释性,表明其特征选择能力更优。
- 当数据足够充分时,SAMM的置信区间接近名义覆盖水平,尽管并非保证为置信区间。
- SAMM的似然函数在计算上难以直接评估,限制了与基于似然方法的直接比较,但模拟结果支持其稳健性。
- 该方法具有通用性,可通过引入额外协变量,扩展至建模其他生物过程,如SNP率或转录因子结合。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。