Skip to main content
QUICK REVIEW

[Paper Review] A proof that a word of length n has less than 1.5n distinct squares

Adrien Thierry|arXiv (Cornell University)|Jan 7, 2020
semigroups and automata theory7 references4 citations
TL;DR

This paper proves that a word of length $ n $ contains fewer than $ 1.5n $ distinct squares, improving upon prior bounds using a refined combinatorial analysis of FS double squares—pairs of squares ending at the same position. By introducing and analyzing $ \alpha $, $ \beta $, $ \epsilon $, and $ \eta $-mates within these structures and leveraging amortized counting via the core of the interrupt, the authors establish a tight bound of $ \frac{3}{2}n $, advancing the long-standing conjecture that the maximum number of distinct squares is less than $ n $.

ABSTRACT

We are interested in the maximal number of distinct squares in a word. This problem was introduced by Fraenkel and Simpson, who presented a bound of 2n for a word of length n, and conjectured that the bound was less than n. Being that the problem is on repetitions, their solution relies on Fine and Wilf's Periodicity lemma. Ilie then refined their result and presented a bound of 2n-O(log n). Lam used an induction to get a bound of 95n/48. Deza, Franek and Thierry achieved a bound of 11n/6 through a combinatorial approach. Using the properties of the core of the interrupt, presented by Thierry, we refined here the combinatorial structures exhibited by Deza, Franek and Thierry to offer a bound of 3n/2.

Motivation & Objective

  • To establish a tighter upper bound on the number of distinct squares in a word of length $ n $, improving upon previous results.
  • To resolve the long-standing conjecture by Fraenkel and Simpson that the number of distinct squares is less than $ n $, by proving a bound of $ \frac{3}{2}n $.
  • To analyze the structural constraints of FS double squares—pairs of squares ending at the same position—using combinatorial tools and mate classifications.
  • To introduce and utilize the concept of the core of the interrupt to redefine and amortize $ \epsilon $-mates, enhancing the induction argument.
  • To provide a complete and exhaustive classification of relations between FS double squares, ensuring the bound is both tight and generalizable.

Proposed method

  • Uses the concept of FS double squares—two distinct squares ending at the same position—to analyze square multiplicities in words.
  • Introduces four types of mates ($ \alpha $, $ \beta $, $ \epsilon $, $ \eta $) to classify and bound the number of overlapping square pairs.
  • Applies Fine and Wilf’s Periodicity Lemma and Crochemore and Rytter’s Three Squares Lemma as foundational tools for periodicity reasoning.
  • Employs an amortized counting strategy based on the core of the interrupt, replacing the earlier inversion factor to better control $ \epsilon $-mates.
  • Develops a recursive induction argument that relies only on the tail of words $ u $ and $ v $, not on their gap, to maintain tight bounds.
  • Uses structural constraints from $ \beta $-mates to limit the number of possible $ \epsilon $-mates, enabling tighter overall bounds.

Experimental results

Research questions

  • RQ1What is the maximal number of distinct squares that can appear in a word of length $ n $?
  • RQ2Can the bound on distinct squares be improved beyond $ 2n $, and if so, by how much?
  • RQ3How do the structural properties of FS double squares—especially their mate relationships—constrain the total number of distinct squares?
  • RQ4Can the amortization of $ \epsilon $-mates using the core of the interrupt lead to a tighter bound than previous methods?
  • RQ5Is the conjecture that the number of distinct squares is less than $ n $ for a word of length $ n $ supported by a bound approaching $ \frac{3}{2}n $?

Key findings

  • The paper proves that a word of length $ n $ contains fewer than $ \frac{3}{2}n $ distinct squares, establishing a new upper bound.
  • The bound is derived by showing that the number of FS double squares is less than $ \frac{n}{2} $, which directly implies the bound on distinct squares.
  • The analysis of $ \epsilon $-mates and their amortization using the core of the interrupt allows the induction to depend only on the tail of words, not their gap.
  • The presence of $ \beta $-mates restricts the number of possible $ \epsilon $-mates, enabling tighter control over square multiplicities.
  • The classification of $ \alpha $, $ \beta $, $ \epsilon $, and $ \eta $-mates provides a complete and exhaustive framework for analyzing overlapping square pairs.
  • The result improves upon previous bounds, including $ \frac{11}{6}n $ by Deza, Franek, and Thierry, and $ \frac{95}{48}n $ by Lam, by achieving $ \frac{3}{2}n $.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.