Skip to main content
QUICK REVIEW

[论文解读] Toward the Combinatorial Limit Theory of Free Words

Danny Rorabaugh|arXiv (Cornell University)|Sep 15, 2015
semigroups and automata theory参考文献 7被引用 3
一句话总结

本文通过分析模式包含与避免,为自由词建立了一套组合极限理论,重点关注不可避免词(如Zimin词)的密度。该研究利用概率方法与结构分析,建立了随机词与长词中期望密度与最小密度的渐近结果,主要贡献包括对Zimin词避免的界限估计,以及基于加倍性质的随机词中密度二分法。

ABSTRACT

Free words are elements of a free monoid, generated over an alphabet via the binary operation of concatenation. Casually speaking, a free word is a finite string of letters. Henceforth, we simply refer to them as words. Motivated by recent advances in the combinatorial limit theory of graphs-notably those involving flag algebras, graph homomorphisms, and graphons-we investigate the extremal and asymptotic theory of pattern containment and avoidance in words. Word V is a factor of word W provided V occurs as consecutive letters within W. W is an instance of V provided there exists a nonerasing monoid homomorphsism ϕ with ϕ(V) = W. For example, using the homomorphism ϕ defined by ϕ(P) = Ror, ϕ(h) = a, and ϕ(D) = baugh, we see that Rorabaugh is an instance of PhD. W avoids V if no factor of W is an instance of V. V is unavoidable provided, over any finite alphabet, there are only finitely many words that avoid V. Unavoidable words were classified by Bean, Ehrenfeucht, and McNulty (1979) and Zimin (1982). We briefly address the following Ramsey-theoretic question: For unavoidable word V and a fixed alphabet, what is the longest a word can be that avoids V? The density of V in W is the proportion of nonempty substrings of W that are instances of V. Since there are 45 substrings in Rorabaugh and 28 of them are instances of PhD, the density of PhD in Rorabaugh is 28/45. We establish a number of asymptotic results for word densities, including the expected density of a word in arbitrarily long, random words and the minimum density of an unavoidable word over arbitrarily long words. This is joint work with Joshua Cooper.

研究动机与目标

  • 将此前仅适用于图的组合极限理论框架,拓展至自由词与模式避免的领域。
  • 研究词密度的极值与渐近行为,特别是Zimin词等不可避免模式的情形。
  • 确定在固定字母表上避免不可避免模式的最长可能词长,解决一个Ramsey理论问题。
  • 刻画在任意长随机词中不可避免词的期望密度与最小密度。
  • 探索自由词极限理论的基础,包括收敛性、极限对象以及准随机性的类比。

提出的方法

  • 使用一阶矩法推导避免Zimin词的词长下界。
  • 应用最小Zimin实例分析,以改进避免界限并理解结构约束。
  • 利用de Bruijn图建模转移概率,并分析长词中子串的分布。
  • 通过区分加倍词与非加倍词,分析随机词中词实例的密度。
  • 应用概率与组合技术,计算期望同态计数与渐近密度。
  • 通过计算验证与Sage代码,验证Zimin词密度与同态计数结果。

实验结果

研究问题

  • RQ1在固定有限字母表上,能避免给定不可避免词V的最长词是什么?
  • RQ2在q个字母的字母表上,均匀随机生成的长词中,固定不可避免词V的期望密度是多少?
  • RQ3在任意长词中,不可避免词V的最小可能密度是多少?
  • RQ4随着词长增加,Zimin词Zn的密度如何渐近变化?
  • RQ5能否基于因子或实例密度,为自由词定义一个富有成效的准随机性概念?

主要发现

  • 本文证明,任何不可避免词V在q个字母表上随机长词中的期望密度收敛至一个明确定义的极限,且对Zimin词推导出了显式界限。
  • 对于Zimin词,所有长词中最小密度与零保持正距离,表明不可避免模式在长字符串中本质上普遍存在。
  • 证明了密度二分法:在随机词中,词V的密度渐近为0或与零保持正距离,具体取决于V是否为加倍词。
  • 随机长度为n的词中,从词V到其同态数的期望值渐近正比于n^k,其中k为V中非重复字母的数量。
  • 计算结果表明,I(Z2, q)与I(Z3, q)(表示随机词中Zimin实例的期望数量)随q可预测增长,计算结果已扩展至q = 6。
  • 本文提供了二元词中避免Z3的最大长度的构造性下界,经计算验证,并给出了一个避免Z4的长二元词。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。