Skip to main content
QUICK REVIEW

[Paper Review] A complexity measure for symbolic sequences and applications to DNA

A. P. Majtey, Ramón Román-Roldán|ArXiv.org|Jun 13, 2006
Fractal and DNA sequence analysis3 citations
TL;DR

This paper proposes a complexity measure for symbolic sequences based on segmenting them into regions of relatively uniform composition and quantifying complexity as the entropy of segment length distribution. The measure satisfies key properties like the one-hump curve, super-additivity, and sensitivity to analysis resolution, and successfully reveals structural differences in DNA sequences, particularly between coding and non-coding regions, with higher complexity in sequences showing long-range correlations and intronic content.

ABSTRACT

We introduce a complexity measure for symbolic sequences. Starting from a segmentation procedure of the sequence, we define its complexity as the entropy of the distribution of lengths of the domains of relatively uniform composition in which the sequence is decomposed. We show that this quantity verifies the properties usually required for a ``good'' complexity measure. In particular it satisfies the one hump property, is super-additive and has the important property of being dependent of the level of detail in which the sequence is analyzed. Finally we apply it to the evaluation of the complexity profile of some genetic sequences.

Motivation & Objective

  • To develop a complexity measure for symbolic sequences that captures both order and disorder, satisfying fundamental properties like the one-hump curve and super-additivity.
  • To investigate how complexity depends on the level of detail in sequence analysis, particularly through adjustable segmentation thresholds.
  • To evaluate the measure’s ability to detect structural and evolutionary features in genomic sequences, such as coding vs. non-coding regions and long-range correlations.
  • To compare the performance of the proposed measure with existing methods, especially in distinguishing biologically meaningful patterns in DNA sequences.

Proposed method

  • Segment a symbolic sequence into contiguous regions of relatively uniform composition using a threshold-based segmentation procedure.
  • Define the complexity as the Shannon entropy of the probability distribution of segment lengths obtained from the segmentation.
  • Adjust the segmentation threshold $ D_u $ to control the level of detail in the analysis, enabling multi-scale complexity evaluation.
  • Apply the measure to real DNA sequences (e.g., human and bacterial genomes) and to computer-generated random sequences for comparison.
  • Use weighted averaging to test the super-additivity property by comparing the complexity of concatenated sequences to the sum of individual complexities.
  • Analyze complexity profiles across different genomic regions (e.g., exons vs. introns) and homologous sequences from different species to assess biological relevance.

Experimental results

Research questions

  • RQ1Does the proposed complexity measure exhibit the one-hump property, indicating maximum complexity at intermediate levels of order and disorder?
  • RQ2How does the complexity value depend on the segmentation threshold $ D_u $, and does this reflect a meaningful multi-scale analysis of sequence structure?
  • RQ3Is the complexity measure super-additive, as required for a robust measure of complexity in composite systems?
  • RQ4Can the complexity measure distinguish between coding and non-coding regions in DNA sequences, particularly those with long-range correlations?
  • RQ5To what extent does the complexity correlate with biological complexity and evolutionary features across homologous sequences of different species?

Key findings

  • The complexity measure exhibits the one-hump property, reaching a maximum at intermediate levels of sequence regularity, confirming its ability to detect complex structures.
  • For a random sequence with the same base composition as the bacterial ECO110k, complexity drops to zero for $ D_u $ in the range 20–50, indicating the method effectively identifies randomness.
  • Human DNA sequences (e.g., HUMTCRADCV) show higher complexity than bacterial sequences across the same $ D_u $ range, reflecting underlying long-range correlations.
  • Complexity profiles do not decay to zero with increasing $ D_u $, unlike previous measures, indicating robustness and sensitivity to structural features.
  • The measure satisfies the super-additivity condition: the complexity of concatenated sequences is greater than or equal to the weighted sum of individual complexities.
  • Complexity values are higher in non-coding regions (introns and intergenic sequences) compared to coding regions, correlating with long-range correlations and evolutionary dynamics.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.