[Paper Review] Coalesced Multi-Output Tsetlin Machines with Clause Sharing
This paper proposes Coalesced Multi-Output Tsetlin Machines with Clause Sharing (CoT-MOTM), a novel architecture that enhances Tsetlin Machines by enabling shared clauses across multiple output labels, improving efficiency and generalization. By coalescing clauses and leveraging pattern memory with confidence-aware clause composition, the method reduces redundancy and accelerates learning while maintaining high accuracy on multi-label classification tasks.
Using finite-state machines to learn patterns, Tsetlin machines (TMs) have obtained competitive accuracy and learning speed across several benchmarks, with frugal memory- and energy footprint. A TM represents patterns as conjunctive clauses in propositional logic (AND-rules), each clause voting for or against a particular output. While efficient for single-output problems, one needs a separate TM per output for multi-output problems. Employing multiple TMs hinders pattern reuse because each TM then operates in a silo. In this paper, we introduce clause sharing, merging multiple TMs into a single one. Each clause is related to each output by using a weight. A positive weight makes the clause vote for output $1$, while a negative weight makes the clause vote for output $0$. The clauses thus coalesce to produce multiple outputs. The resulting coalesced Tsetlin Machine (CoTM) simultaneously learns both the weights and the composition of each clause by employing interacting Stochastic Searching on the Line (SSL) and Tsetlin Automata (TA) teams. Our empirical results on MNIST, Fashion-MNIST, and Kuzushiji-MNIST show that CoTM obtains significantly higher accuracy than TM on $50$- to $1$K-clause configurations, indicating an ability to repurpose clauses. E.g., accuracy goes from $71.99$% to $89.66$% on Fashion-MNIST when employing $50$ clauses per class (22 Kb memory). While TM and CoTM accuracy is similar when using more than $1$K clauses per class, CoTM reaches peak accuracy $3 imes$ faster on MNIST with $8$K clauses. We further investigate robustness towards imbalanced training data. Our evaluations on imbalanced versions of IMDb- and CIFAR10 data show that CoTM is robust towards high degrees of class imbalance. Being able to share clauses, we believe CoTM will enable new TM application domains that involve multiple outputs, such as learning language models and auto-encoding.
Motivation & Objective
- To address the inefficiency and redundancy in multi-output Tsetlin Machine designs by enabling shared clause usage across multiple output labels.
- To improve model generalization and training efficiency through a coalesced clause structure that reduces parameter proliferation.
- To introduce a pattern memory mechanism that encodes clause composition and memorization confidence using values from {1, 2, ..., 2N}.
- To enable scalable multi-label classification by mapping input patterns to clauses with confidence-aware representations.
- To reduce computational overhead in multi-output learning by reusing clauses across different output heads.
Proposed method
- The method employs a coalesced clause structure where clauses are shared across multiple output labels, minimizing redundant clause creation.
- It uses a matrix representation for pattern memory, with rows denoting clauses and columns representing input variables or their negations (xk or ¬xk).
- Each matrix entry takes values from {1, 2, ..., 2N}, encoding both clause composition and memorization confidence.
- Input patterns are mapped to clauses via a two-action mapping (a ∈ {0, 1}) to determine clause activation based on variable presence or negation.
- The architecture supports multi-output learning by assigning shared clauses to multiple output heads, reducing model complexity.
- Clause sharing is enforced through a unified pattern memory that tracks clause composition and confidence levels across all outputs.
Experimental results
Research questions
- RQ1How can clause sharing across multiple output labels improve efficiency in multi-output Tsetlin Machines?
- RQ2To what extent does coalescing clauses reduce redundancy and parameter count in multi-label learning?
- RQ3Can confidence-aware clause representation enhance memorization and generalization in Tsetlin Machines?
- RQ4How does shared clause usage affect classification accuracy and convergence speed?
- RQ5What is the impact of pattern memory structure on scalability and performance in multi-output settings?
Key findings
- Clause sharing significantly reduces the number of required clauses, decreasing model complexity and training overhead.
- The confidence-aware pattern memory representation enables stronger memorization of relevant patterns, improving learning efficiency.
- Coalesced clause structures maintain high classification accuracy while reducing redundancy across multiple output labels.
- The method demonstrates improved scalability in multi-label classification tasks due to shared clause utilization.
- Mapping inputs to clauses via a two-action mechanism (a ∈ {0,1}) enables effective and efficient pattern encoding.
- The architecture achieves competitive performance with fewer parameters and faster convergence compared to non-shared counterparts.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.