[Paper Review] Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?
This paper proposes reframing large language models (LLMs) as complex systems, advocating for a top-down behavioral taxonomy to guide mechanistic interpretability and future research. By categorizing emergent behaviors across tasks, the framework enables systematic analysis of LLMs independent of architecture, offering a foundation for replicable, theory-driven research despite model opacity and rapid architectural evolution.
Coaxing out desired behavior from pretrained models, while avoiding undesirable ones, has redefined NLP and is reshaping how we interact with computers. What was once a scientific engineering discipline-in which building blocks are stacked one on top of the other-is arguably already a complex systems science, in which emergent behaviors are sought out to support previously unimagined use cases. Despite the ever increasing number of benchmarks that measure task performance, we lack explanations of what behaviors language models exhibit that allow them to complete these tasks in the first place. We argue for a systematic effort to decompose language model behavior into categories that explain cross-task performance, to guide mechanistic explanations and help future-proof analytic research.
Motivation & Objective
- Address the lack of systematic understanding of emergent behaviors in large language models (LLMs), which currently limits interpretability and theory-building.
- Overcome the limitations of benchmark-centric evaluation by shifting focus from performance metrics to identifying and categorizing high-level model behaviors.
- Develop a top-down behavioral taxonomy to guide bottom-up mechanistic analysis, ensuring research targets meaningful phenomena rather than arbitrary model components.
- Establish a metamodel framework that predicts regularities in LLM outputs, enabling more robust and generalizable explanations of model behavior.
- Promote open-source model availability as essential for replicable, long-term scientific study of generative models, especially given the dominance of proprietary systems.
Proposed method
- Introduce a thought experiment involving the 'Newformer'—a hypothetical, non-Transformer, state-of-the-art generative model—to highlight the limitations of current model-interpretation approaches.
- Propose a hierarchical, top-down taxonomy of LLM behaviors (e.g., paraphrasing, repetition, in-context learning) to serve as a functional framework for guiding mechanistic investigations.
- Use existing interpretability findings (e.g., induction heads, copying heads) as evidence that behavior-based categorization leads to deeper mechanistic insights.
- Frame generative models as complex systems due to their emergent, non-engineered behaviors, drawing parallels to natural complex systems like biological or chemical systems.
- Argue that simulation and repeatability in LLMs offer advantages over physical complex systems, enabling controlled, observer-free experimentation.
- Stress the necessity of open-source models to sustain long-term, replicable research, especially when proprietary models dominate capabilities.

Experimental results
Research questions
- RQ1How can we systematically categorize the emergent behaviors of large language models to guide mechanistic interpretation?
- RQ2What role does a top-down behavioral taxonomy play in improving the efficiency and relevance of bottom-up mechanistic analysis?
- RQ3Why do benchmarks fail to capture or explain the behaviors that underlie LLM performance across tasks?
- RQ4How can we develop a metamodel that predicts regularities in LLM outputs without relying on architectural specifics?
- RQ5In what ways do generative models resemble complex systems, and how does this perspective shift the goals of NLP research?
Key findings
- Emergent behaviors in LLMs—such as in-context learning and phrase repetition—are not engineered but discovered through training, indicating complex system dynamics.
- The absence of a shared behavioral vocabulary hinders progress, even when models are open-sourced, as seen in the difficulty of explaining behaviors like anisotropic embedding spaces.
- Existing interpretability work (e.g., induction heads) demonstrates that behavior-based categorization leads to actionable mechanistic insights, validating the proposed taxonomy approach.
- Despite architectural differences, shared high-level behaviors (e.g., paraphrasing, copying) can be used to compare models like the fictional Newformer and Transformers, enabling cross-model analysis.
- Open-source models are essential for replicable, long-term scientific study, as proprietary models limit access to the behaviors that drive performance.
- Generative models are more amenable to scientific study than many natural complex systems due to their simulability, repeatability, and lack of observer effects in controlled experiments.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.