[论文解读] Multiple Sequence Alignment is not a Solved Problem
本文认为,多重序列比对(MSA)仍然是一个未解决的计算问题,因为它旨在推断同源性——一种历史上定义但不可观测的事件——而非优化可度量的目标函数。作者提出,应将MSA重新框架为一种整合多种生物学标准(组成、拓扑、功能、个体发育及一致性)的推断问题,而非仅基于相似性的优化任务。
Multiple sequence alignment is a basic procedure in molecular biology, and it is often treated as being essentially a solved computational problem. However, this is not so, and here I review the evidence for this claim, and outline the requirements for a solution. The goal of alignment is often stated to be to juxtapose nucleotides (or their derivatives, such as amino acids) that have been inherited from a common ancestral nucleotide (although other goals are also possible). Unfortunately, this is not an operational definition, because homology (in this sense) refers to unique and unobservable historical events, and so there can be no objective mathematical function to optimize. Consequently, almost all algorithms developed for multiple sequence alignment are based on optimizing some sort of compositional similarity (similarity = homology + analogy). As a result, many, if not most, practitioners either manually modify computer-produced alignments or they perform de novo manual alignment, especially in the field of phylogenetics. So, if homology is the goal, then multiple sequence alignment is not yet a solved computational problem. Several criteria have been developed by biologists to help them identify potential homologies (compositional, ontogenetic, topographical and functional similarity, plus conjunction and congruence), and these criteria can be applied to molecular data, in principle. Current computer programs do implement one (or occasionally two) of these criteria, but no program implements them all. What is needed is a program that evaluates all of the evidence for the sequence homologies, optimizes their combination, and thus produces the best hypotheses of homology. This is basically an inference problem not an optimization problem.
研究动机与目标
- 挑战分子生物学中广泛存在的‘多重序列比对是已解决的计算问题’的假设。
- 强调同源性——MSA的真正目标——是不可观测的历史事件,因此无法直接优化。
- 论证当前的MSA算法依赖于优化组成相似性(同源性+类比性),导致生物学上不可靠的结果。
- 倡导从基于优化的MSA转向基于推断的框架,以评估所有可用的生物学证据来推断同源性。
- 呼吁开发一种计算方法,整合多种标准(组成、拓扑、功能、个体发育、一致性),以生成最合理的序列同源性假设。
提出的方法
- 将MSA重新框架为假设检验的推断问题,而非优化问题。
- 识别并整合同源性推断的五个关键生物学标准:组成相似性、拓扑相似性、功能相似性、个体发育相似性以及独立数据集间的一致性。
- 提出目前尚无程序同时实现全部五个标准。
- 强调同源性无法通过单一数学函数推导,因为它依赖于历史性的、不可观测的事件。
- 倡导一种计算框架,以评估并整合所有五个标准的证据,从而推断最合理的同源性假设。
- 强调此类系统不会优化单一得分,而是权衡并综合多样化的生物学证据。
实验结果
研究问题
- RQ1尽管计算工具被广泛使用,为何多重序列比对仍不被视为已解决的问题?
- RQ2当前依赖于优化组成相似性的MSA算法的根本局限性是什么?
- RQ3当同源性(即共同祖先关系)是一种不可观测的历史事件时,如何可靠地推断它?
- RQ4除了序列相似性外,还可以使用哪些生物学标准来识别潜在的同源序列?
- RQ5一种能够整合多种标准以推断同源性而非优化单一得分的计算框架会是什么样子?
主要发现
- 多重序列比对并非已解决的问题,因为其目标——同源性——是不可观测的历史事件,无法直接优化。
- 当前的MSA算法优化的是组成相似性,这混淆了同源性与类比性,导致生物学上误导的结果。
- 许多从业者手动修改或重新进行比对,尤其是在系统发育分析中,表明自动化工具存在局限性。
- 目前尚无程序能同时实现全部五个标准(组成、拓扑、功能、个体发育及一致性)进行同源性推断。
- 核心问题并非算法复杂性,而是概念性的:MSA并非优化问题,而是需要整合多样化生物学证据的推断问题。
- 需要一种新的计算框架,以评估并整合所有可用的生物学证据,从而生成最合理的同源性假设。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。