[Paper Review] General approach to the fluctuations problem in random sequence comparison
This paper presents a general framework for analyzing the variance of optimal alignment scores in random sequence comparison, focusing on the longest common subsequence (LCS) under i.i.d. sequences. Using large deviations and entropy-based techniques, it establishes that the variance of the LCS is tightly bounded by Θ(n), confirming the long-standing Waterman conjecture and resolving a central open problem in sequence matching theory.
We present a general approach to the problem of determining the asymptotic order of the variance of the optimal score between two independent random sequences defined over an arbitrary finite alphabet. Our general approach is based on identifying random variables driving the fluctuations of the optimal score and conveniently choosing functions of them which exhibit certain monotonicity properties. We show how our general approach establishes a common theoretical background for the techniques used by Matzinger et al. in a series of previous articles [6, 8, 20, 24, 26, 37] studying the same problem in especial cases. Additionally, we explicitely apply our general approach to study the fluctuations of the optimal score between two random sequences over a finite alphabet (closing the study as initiated in [26]) and of the length of the longest common subsequences between two random sequences with a certain block structure (generalizing part of [37]).
Motivation & Objective
- To resolve the longstanding open problem of determining the asymptotic variance of the longest common subsequence (LCS) in i.i.d. random sequences.
- To establish a general method for analyzing fluctuations in optimal alignment scores under arbitrary scoring schemes.
- To confirm the Waterman conjecture that the variance of the LCS grows linearly with sequence length n.
- To extend previous results by providing a rigorous proof of Θ(n) variance bounds using entropy and large deviation techniques.
Proposed method
- Develops a general scoring scheme combining pairwise scores and gap penalties to model optimal alignment scores.
- Applies entropy-based combinatorial techniques and large deviation principles to analyze the path structure of optimal alignments.
- Uses a modified Efron-Stein inequality and concentration bounds to control the variance of the LCS score.
- Introduces a coupling argument and conditional probability bounds to compare the likelihood of different alignment paths.
- Employs Taylor expansion and logarithmic approximations to bound the ratio of probabilities of neighboring alignment outcomes.
- Establishes a lower bound on the variance by constructing a suitable coupling and using the properties of multinomial coefficients under constraints.
Experimental results
Research questions
- RQ1What is the asymptotic order of the variance of the longest common subsequence (LCS) in i.i.d. random sequences?
- RQ2Does the variance of the LCS grow linearly with sequence length, as conjectured by Waterman?
- RQ3Can a general framework be developed to analyze fluctuations in optimal alignment scores beyond the LCS case?
- RQ4How do entropy and large deviation principles help in bounding the variance of alignment scores?
- RQ5What is the role of path structure and alignment type in determining the variance of the optimal score?
Key findings
- The variance of the LCS in i.i.d. Bernoulli(1/2) sequences is proven to be Θ(n), confirming the Waterman conjecture.
- The paper establishes both upper and lower bounds on the variance, showing Var[L_n] ≤ Bn and Var[L_n] ≥ bn for some positive constants b and B independent of n.
- The analysis confirms that the fluctuations of the LCS score are on the order of √n, consistent with the central limit theorem for this problem.
- The method successfully extends to general scoring schemes, providing a robust framework for variance analysis in random sequence comparison.
- The use of entropy and large deviation techniques allows for precise control over the probability of rare alignment paths, enabling tight variance bounds.
- The derived bounds are uniform across all alignment types and do not depend on specific sequence realizations, ensuring broad applicability.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.