Skip to main content
QUICK REVIEW

[論文レビュー] Deciding the Confusability of Words under Tandem Repeats

Yeow Meng Chee, Johan Chrisnata|arXiv (Cornell University)|Jul 13, 2017
Algorithms and Data Compression参考文献 17被引用数 11
ひとこと要約

本稿では、k ≤ 3 の場合に限って、文字列の並び替えと重複挿入に基づく confusability を線形時間で決定するアルゴリズムを提示している。この手法は、文字列の置換とラベル系による新しい confusability の特徴付けに基づいている。主な貢献は、k ≤ 3 に対する完全な決定手続きであり、最適な並び替え重複コードのサイズに関する、改善された上界と下界を提示しており、長さ 20 までについて正確な値を算出し、長さ 21–30 については再帰的構成法を用いている。

ABSTRACT

Tandem duplication in DNA is the process of inserting a copy of a segment of DNA adjacent to the original position. Motivated by applications that store data in living organisms, Jain {\em et al.} (2016) proposed the study of codes that correct tandem duplications to improve the reliability of data storage. We investigate algorithms associated with the study of these codes. Two words are said to be ${\le}k$-confusable if there exists two sequences of tandem duplications of lengths at most $k$ such that the resulting words are equal. We demonstrate that the problem of deciding whether two words is ${\le}k$-confusable is linear-time solvable through a characterisation that can be checked efficiently for $k=3$. Combining with previous results, the decision problem is linear-time solvable for $k\le 3$. We conjecture that this problem is undecidable for $k>3$. Using insights gained from the algorithm, we study the size of tandem-duplication codes. We improve the previous known upper bound and then construct codes with larger sizes as compared to the previous constructions. We determine the sizes of optimal tandem-duplication codes for lengths up to twenty, develop recursive methods to construct tandem-duplication codes for all word lengths, and compute explicit lower bounds for the size of optimal tandem-duplication codes for lengths from 21 to 30.

研究の動機と目的

  • 並び替え重複の長さが最大 k である場合に、2つの語がconfusableかどうかを効率的に判定するアルゴリズムの開発。
  • k ≤ 3 の場合に、文字列の置換とラベル系を用いたconfusability問題の特徴付けを行い、線形時間での決定を可能にする。
  • 特に短い語の長さにおいて、最適な並び替え重複コードのサイズに関する上界と下界の改善。
  • 任意の長さに対して大きな並び替え重複コードを生成する再帰的メソッドの構築。
  • 長さ 21 から 30 までの最適コードの正確で明示的な下界の計算。

提案手法

  • 並び替え重複の下での構造的不変量を追跡するためのラベル系を導入し、効率的なconfusabilityのチェックを可能にする。
  • 並び替え重複規則の下での簡約形の同値性に基づく、≤k-confusability の特徴付けを構築する。
  • 動的計画法と MaxCliqueDyn を用いて、長さ 20 までの一連の最適コードのサイズを計算する。
  • 不可約語の構造とその接尾語に依存する再帰的構成法(命題 18)を提案する。
  • 子孫の錐の交差と形式的言語理論を用いて、特に k > 3 の場合のconfusabilityを分析する。
  • 情報理論における表現力と容量の概念を用いて、コードの性能とレートを評価する。

実験結果

リサーチクエスチョン

  • RQ1k ≤ 3 の場合に、長さが最大 k の並び替え重複の下で、2つの語のconfusabilityを線形時間で決定可能か?
  • RQ2長さが最大 k の並び替え重複を是正できるコードの最大サイズは何か? そして、このようなコードは効率的に構築可能か?
  • RQ3コードサイズの上界と下界はどのように比較されるか? 特に短い語の長さにおいて、より鋭い境界を確立できるか?
  • RQ4k > 3 の場合に、confusability 問題は決定可能か? また、線形時間解法が不可能となる構造的性質は何か?
  • RQ5任意の語の長さに対して、大きな並び替え重複コードを生成する再帰的構成法を導出可能か?

主な発見

  • ≤k-confusability を決定する問題は、k ≤ 3 の場合に、ラベル系と簡約形に基づく特徴付けにより線形時間で解ける。
  • 長さ 1 から 20 までの最適並び替え重複コードのサイズについて、正確な値が特定された。
  • 長さ 21 から 30 については、再帰的構成法を用いて明示的な下界が計算され、長さ 30 で最大 2,464,419 のコードサイズに達している。
  • 新しい構成法により、長さ 21–30 において、従来の構成法よりも最大 6.74% のレート向上が達成された。
  • 命題 4 からの上界は、長さ ≥11 において、式 (1) からの境界よりも鋭く、新しい構成法は従来の下界を上回っている。
  • 本稿では、k ≥ 4 の場合に、構造的不変量の喪失と文脈自由言語における交差空集合問題の決定不能性のため、confusability 問題が決定不能であると予想している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。