[Paper Review] Toward the Combinatorial Limit Theory of Free Words
This paper develops a combinatorial limit theory for free words by analyzing pattern containment and avoidance, focusing on densities of unavoidable words like Zimin words. It establishes asymptotic results for expected and minimum densities in random and long words, using probabilistic methods and structural analysis, with key contributions including bounds on Zimin word avoidance and a density dichotomy in random words based on doubling properties.
Free words are elements of a free monoid, generated over an alphabet via the binary operation of concatenation. Casually speaking, a free word is a finite string of letters. Henceforth, we simply refer to them as words. Motivated by recent advances in the combinatorial limit theory of graphs-notably those involving flag algebras, graph homomorphisms, and graphons-we investigate the extremal and asymptotic theory of pattern containment and avoidance in words. Word V is a factor of word W provided V occurs as consecutive letters within W. W is an instance of V provided there exists a nonerasing monoid homomorphsism ϕ with ϕ(V) = W. For example, using the homomorphism ϕ defined by ϕ(P) = Ror, ϕ(h) = a, and ϕ(D) = baugh, we see that Rorabaugh is an instance of PhD. W avoids V if no factor of W is an instance of V. V is unavoidable provided, over any finite alphabet, there are only finitely many words that avoid V. Unavoidable words were classified by Bean, Ehrenfeucht, and McNulty (1979) and Zimin (1982). We briefly address the following Ramsey-theoretic question: For unavoidable word V and a fixed alphabet, what is the longest a word can be that avoids V? The density of V in W is the proportion of nonempty substrings of W that are instances of V. Since there are 45 substrings in Rorabaugh and 28 of them are instances of PhD, the density of PhD in Rorabaugh is 28/45. We establish a number of asymptotic results for word densities, including the expected density of a word in arbitrarily long, random words and the minimum density of an unavoidable word over arbitrarily long words. This is joint work with Joshua Cooper.
Motivation & Objective
- To extend the framework of combinatorial limit theory—previously developed for graphs—to the domain of free words and pattern avoidance.
- To investigate the extremal and asymptotic behavior of word densities, particularly for unavoidable patterns such as Zimin words.
- To determine the longest possible words that avoid unavoidable patterns over fixed alphabets, addressing a Ramsey-theoretic question.
- To characterize the expected and minimum densities of unavoidable words in arbitrarily long random words.
- To explore the foundations of a limit theory for free words, including convergence, limit objects, and quasirandomness analogues.
Proposed method
- Uses the first moment method to derive lower bounds on the length of words avoiding Zimin words.
- Applies minimal Zimin-instance analysis to refine avoidance bounds and understand structural constraints.
- Employs the de Bruijn graph to model transition probabilities and analyze substring distributions in long words.
- Analyzes the density of word instances in random words by distinguishing between doubled and nondoubled words.
- Applies probabilistic and combinatorial techniques to compute expected homomorphism counts and asymptotic densities.
- Uses computational verification and Sage code to validate results on Zimin word densities and homomorphism counts.
Experimental results
Research questions
- RQ1What is the longest word that can avoid a given unavoidable word V over a fixed finite alphabet?
- RQ2What is the expected density of a fixed unavoidable word V in a uniformly random long word over a q-letter alphabet?
- RQ3What is the minimum possible density of an unavoidable word V over arbitrarily long words?
- RQ4How do the densities of Zimin words Zn behave asymptotically as word length increases?
- RQ5Can a productive notion of quasirandomness be defined for free words based on factor or instance densities?
Key findings
- The paper establishes that the expected density of any unavoidable word V in a random long word over a q-letter alphabet converges to a well-defined limit, with explicit bounds derived for Zimin words.
- For Zimin words, the minimum density over all long words is bounded away from zero, indicating that unavoidable patterns are inherently prevalent in long strings.
- A density dichotomy is proven: for random words, the density of a word V is asymptotically either 0 or bounded away from zero, depending on whether V is doubled or not.
- The expected number of homomorphisms from a word V to a random word of length n is asymptotically proportional to n^k, where k is the number of nonrecurring letters in V.
- Computational results show that I(Z2, q) and I(Z3, q), representing the expected number of Zimin instances in random words, grow predictably with q, with values computed up to q = 6.
- The paper provides a constructive lower bound on the maximum length of binary words avoiding Z3, verified computationally, and presents a long binary word avoiding Z4.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.