Skip to main content
QUICK REVIEW

[论文解读] Governance Architecture for Neural Network Superposition: A Structural Solution to Hallucination via Routing and Interference Filtering

Nelson Elhage, Tristan Hume|arXiv (Cornell University)|Jan 1, 2022
Model Reduction and Neural Networks被引用 37
一句话总结

本论文通过 toy 模型的 superposition 研究神经网络中的 polysemanticity,揭示相变、与 uniform polytopes 的几何联系,以及与 adversarial examples 的关联,并对 mechanistic interpretability 的含义。

ABSTRACT

Neural networks often pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity' which makes interpretability much more challenging. This paper provides a toy model where polysemanticity can be fully understood, arising as a result of models storing additional sparse features in "superposition." We demonstrate the existence of a phase change, a surprising connection to the geometry of uniform polytopes, and evidence of a link to adversarial examples. We also discuss potential implications for mechanistic interpretability.

研究动机与目标

  • Explain polysemanticity as a consequence of storing sparse features in superposition.
  • Characterize the conditions under which superposition emerges as a phase change.
  • Link the geometry of superposition to uniform polytopes and adversarial examples.
  • Discuss implications for mechanistic interpretability and model governance.

提出的方法

  • Introduce toy models that realize feature superposition in neurons.
  • Analyze phase transitions associated with superposition.
  • Draw connections between superposition geometry and uniform polytopes.
  • Investigate relationships to adversarial examples.
  • Discuss interpretability and governance implications for neural networks.

实验结果

研究问题

  • RQ1What causes polysemanticity to arise in neural networks as a form of feature superposition?
  • RQ2Under what conditions does a phase change occur leading to superposition behavior?
  • RQ3How does the geometry of superposition relate to uniform polytopes and adversarial vulnerability?
  • RQ4What are the implications of superposition for mechanistic interpretability and model governance?

主要发现

  • Evidence of a phase change leading to superposition in toy models.
  • Identification of a surprising link between superposition geometry and uniform polytopes.
  • Indicators suggesting a relation between superposition and adversarial examples.
  • Discussion of how routing and interference filtering could govern superposition.
  • Implications for interpretability through a structural viewpoint on neuron-level representations.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。