Skip to main content
QUICK REVIEW

[Paper Review] Governance Architecture for Neural Network Superposition: A Structural Solution to Hallucination via Routing and Interference Filtering

Nelson Elhage, Tristan Hume|arXiv (Cornell University)|Jan 1, 2022
Model Reduction and Neural Networks37 citations
TL;DR

The paper studies polysemanticity in neural networks through toy models of superposition, revealing a phase change, geometric connections to uniform polytopes, and links to adversarial examples, with implications for mechanistic interpretability.

ABSTRACT

Neural networks often pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity' which makes interpretability much more challenging. This paper provides a toy model where polysemanticity can be fully understood, arising as a result of models storing additional sparse features in "superposition." We demonstrate the existence of a phase change, a surprising connection to the geometry of uniform polytopes, and evidence of a link to adversarial examples. We also discuss potential implications for mechanistic interpretability.

Motivation & Objective

  • Explain polysemanticity as a consequence of storing sparse features in superposition.
  • Characterize the conditions under which superposition emerges as a phase change.
  • Link the geometry of superposition to uniform polytopes and adversarial examples.
  • Discuss implications for mechanistic interpretability and model governance.

Proposed method

  • Introduce toy models that realize feature superposition in neurons.
  • Analyze phase transitions associated with superposition.
  • Draw connections between superposition geometry and uniform polytopes.
  • Investigate relationships to adversarial examples.
  • Discuss interpretability and governance implications for neural networks.

Experimental results

Research questions

  • RQ1What causes polysemanticity to arise in neural networks as a form of feature superposition?
  • RQ2Under what conditions does a phase change occur leading to superposition behavior?
  • RQ3How does the geometry of superposition relate to uniform polytopes and adversarial vulnerability?
  • RQ4What are the implications of superposition for mechanistic interpretability and model governance?

Key findings

  • Evidence of a phase change leading to superposition in toy models.
  • Identification of a surprising link between superposition geometry and uniform polytopes.
  • Indicators suggesting a relation between superposition and adversarial examples.
  • Discussion of how routing and interference filtering could govern superposition.
  • Implications for interpretability through a structural viewpoint on neuron-level representations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.