Skip to main content
QUICK REVIEW

[論文レビュー] Toward the Combinatorial Limit Theory of Free Words

Danny Rorabaugh|arXiv (Cornell University)|Sep 15, 2015
semigroups and automata theory参考文献 7被引用数 3
ひとこと要約

この論文は、パターンの包含と回避を分析することにより、自由語に対する組合せ的極限理論を構築し、Zimin語のような避けられない語の密度に注目する。確率的技法と構造的解析を用いて、ランダム語および長い語における期待密度および最小密度に関する漸近的結果を確立し、主な貢献としてZimin語の回避に関する境界と、二重化性に基づくランダム語における密度の二分法が得られる。

ABSTRACT

Free words are elements of a free monoid, generated over an alphabet via the binary operation of concatenation. Casually speaking, a free word is a finite string of letters. Henceforth, we simply refer to them as words. Motivated by recent advances in the combinatorial limit theory of graphs-notably those involving flag algebras, graph homomorphisms, and graphons-we investigate the extremal and asymptotic theory of pattern containment and avoidance in words. Word V is a factor of word W provided V occurs as consecutive letters within W. W is an instance of V provided there exists a nonerasing monoid homomorphsism ϕ with ϕ(V) = W. For example, using the homomorphism ϕ defined by ϕ(P) = Ror, ϕ(h) = a, and ϕ(D) = baugh, we see that Rorabaugh is an instance of PhD. W avoids V if no factor of W is an instance of V. V is unavoidable provided, over any finite alphabet, there are only finitely many words that avoid V. Unavoidable words were classified by Bean, Ehrenfeucht, and McNulty (1979) and Zimin (1982). We briefly address the following Ramsey-theoretic question: For unavoidable word V and a fixed alphabet, what is the longest a word can be that avoids V? The density of V in W is the proportion of nonempty substrings of W that are instances of V. Since there are 45 substrings in Rorabaugh and 28 of them are instances of PhD, the density of PhD in Rorabaugh is 28/45. We establish a number of asymptotic results for word densities, including the expected density of a word in arbitrarily long, random words and the minimum density of an unavoidable word over arbitrarily long words. This is joint work with Joshua Cooper.

研究の動機と目的

  • グラフの分野で既に開発された組合せ的極限理論の枠組みを、自由語とパターン回避の分野へ拡張すること。
  • 特にZimin語のような避けられないパターンに対する語の密度の極値的および漸近的挙動を調査すること。
  • 固定アルファベット上での避けられないパターンを回避することができる語の最大長を特定することにより、ラマヌジャン的問題に取り組むこと。
  • 任意に長いランダム語における避けられない語の期待密度および最小密度を特定すること。
  • 自由語のための極限理論の基礎を探究すること、収束、極限対象、および擬似ランダム性の類似概念を含むこと。

提案手法

  • Zimin語を回避する語の長さに関する下界を導出するために、第一モーメント法を用いる。
  • 最小Ziminインスタンス解析を適用して、回避に関する境界を精緻化し、構造的制約を理解する。
  • de Bruijnグラフを用いて遷移確率をモデル化し、長い語における部分語の分布を分析する。
  • 語のインスタンスの密度を、二重化された語とそうでない語に区別することで、ランダム語における密度を分析する。
  • 確率的および組合せ的技法を用いて、期待ホモモーティズム数と漸近的密度を計算する。
  • 計算的検証とSageコードを用いて、Zimin語の密度およびホモモーティズム数に関する結果を検証する。

実験結果

リサーチクエスチョン

  • RQ1固定された有限アルファベット上での与えられた避けられない語 V を回避できる語の最大長は何か?
  • RQ2q 文字のアルファベット上での一様ランダムな長い語における、固定された避けられない語 V の期待密度は何か?
  • RQ3任意に長い語上での避けられない語 V の最小密度は最小でどれほどか?
  • RQ4語の長さが増加する際、Zimin語 Zn の密度は漸近的にどのように振る舞うか?
  • RQ5因子またはインスタンスの密度に基づいて、自由語に対して擬似ランダム性の生産的な概念を定義できるか?

主な発見

  • 本論文は、q 文字のアルファベット上でのランダムな長い語における避けられない語 V の期待密度が、明示的な境界が得られる well-defined な極限に収束することを確立した。
  • Zimin語に対しては、すべての長い語における最小密度がゼロから離れていることが示され、避けられないパターンが長い文字列において本質的に顕在することを示している。
  • 密度の二分法が証明された:ランダム語において、語 V の密度は、V が二重化されているかどうかに応じて、漸近的に 0 またはゼロから離れている。
  • 語 V から長さ n のランダム語へのホモモーティズムの期待数は、漸近的に V の再発しない文字の数 k に比例して n^k に比例する。
  • 計算結果から、I(Z2, q) および I(Z3, q) —— それぞれランダム語におけるZiminインスタンスの期待数を表す —— は q とともに予測可能な速度で増加し、q = 6 まで計算された。
  • 本論文は、Z3 を回避する2文字語の最大長に対する構成的下界を提示し、計算的に検証した。また、Z4 を回避する長い2文字語を提示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。