Skip to main content
QUICK REVIEW

[Paper Review] Bounds on Zimin Word Avoidance

Joshua Cooper, Danny Rorabaugh|arXiv (Cornell University)|Sep 10, 2014
Language, Linguistics, Cultural Analysis3 references8 citations
TL;DR

This paper establishes tight bounds on the maximum length of words that avoid unavoidable Zimin words over finite alphabets. Using induction and the pigeonhole principle, it proves an upper bound of $ f(n,q) \leq (2q+1)^{n-1} $, while applying the first moment method to derive a matching lower bound of $ f(n,q) \geq q^{2(n-1)(1+o(1))} $, closing the gap between known upper and lower estimates for word avoidance lengths in combinatorics on words.

ABSTRACT

A pattern is encountered in a word if some infix of the word is the image of the pattern under some non-erasing morphism. A pattern p is unavoidable if, over every finite alphabet, every sufficiently long word encounters p. A theorem by Zimin and independently by Bean, Ehrenfeucht and McNulty states that a pattern over n distinct variables is unavoidable if, and only if, p itself is encountered in the n-th Zimin pattern. Given an alphabet size k, we study the minimal length f(n,k) such that every word of length f(n,k) encounters the n-th Zimin pattern. It is known that f is upper-bounded by a tower of exponentials. Our main result states that f(n,k) is lower-bounded by a tower of n-3 exponentials, even for k=2. To the best of our knowledge, this improves upon a previously best-known doubly-exponential lower bound. As a further result, we prove a doubly-exponential upper bound for encountering Zimin patterns in the abelian sense.

Motivation & Objective

  • To determine the maximum length of words that avoid unavoidable Zimin words over finite alphabets.
  • To close the gap between known upper and lower bounds on the function $ f(n,q) $, which gives the minimal length $ M $ such that every $ q $-ary word of length $ M $ contains $ Z_n $ as a subword.
  • To apply Ramsey-theoretic principles to word avoidance, extending classical results on unavoidable words in combinatorics on words.
  • To refine existing bounds using the first moment method and structural induction, improving upon prior Ackermann-type bounds.

Proposed method

  • Uses induction on $ n $ to prove an upper bound $ f(n,q) \leq (2q+1)^{n-1} $, relying on the pigeonhole principle to force repeated instances of $ Z_n $ in long words.
  • Applies the first moment method to a random word model, computing the expected number of $ Z_n $-instances to derive a lower bound on $ f(n,q) $.
  • Employs a recursive decomposition of $ Z_{n+1} $-instances as $ UVU $, where $ U $ is a $ Z_n $-instance, to overcount and bound the number of such instances.
  • Uses monotonicity of instance probabilities with word length to compare $ |C(n,q,M)| $ across lengths and derive concentration bounds.
  • Establishes that $ |C(n,q,M)| \leq \left(\frac{q}{q-1}\right)^{n-1} q^{M - 2^n + n + 1} $, which is crucial for bounding the probability of avoiding $ Z_n $.
  • Leverages the fact that $ Z_n $ has length $ 2^n - 1 $, and uses the 2-adic order characterization to define $ Z_n $ recursively.

Experimental results

Research questions

  • RQ1What is the maximum length of a word over a $ q $-letter alphabet that avoids the $ n $th Zimin word $ Z_n $?
  • RQ2How do the upper and lower bounds on $ f(n,q) $, the minimal length guaranteeing $ Z_n $-occurrence, compare, and can they be tightened?
  • RQ3Can probabilistic methods such as the first moment method yield a nontrivial lower bound on $ f(n,q) $ that matches the known upper bound asymptotically?
  • RQ4How does the structure of $ Z_n $-instances constrain the growth of avoidable words over finite alphabets?
  • RQ5Can the bound $ f(n,q) \leq (2q+1)^{n-1} $ be improved, or is it tight up to lower-order terms?

Key findings

  • The paper establishes an upper bound $ f(n,q) \leq (2q+1)^{n-1} $, which is primitive recursive and significantly tighter than prior Ackermann-type bounds.
  • A lower bound of $ f(n,q) \geq q^{2(n-1)(1+o(1))} $ is derived using the first moment method, showing that avoidance is impossible beyond this length.
  • The bounds are asymptotically tight: the lower bound matches the upper bound up to a $ (1+o(1)) $ factor in the exponent, closing the gap between known estimates.
  • For $ q=2 $, the paper computes $ f(3,2) = 29 $ and $ f(4,2) \geq 10483 $, providing concrete values for small $ n $.
  • The probability that a random word of length $ M $ is a $ Z_n $-instance is bounded above by $ \left(\frac{q}{q-1}\right)^{n-1} q^{-2^n + n + 1} $, which decays exponentially with $ n $.
  • The monotonicity of instance probability with word length ensures that if $ E[X] < 1 $, then there exists at least one word of length $ M $ avoiding $ Z_n $, enabling the lower bound derivation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.