[论文解读] Disentangling Syntax and Semantics in the Brain with Deep Networks
本文提出一种分类法,将 GPT-2 的激活因子分解为词汇、组合、句法和语义表示,然后将这些分解后的激活映射到 345 名参与叙事的受试者的 fMRI 大脑数据,揭示分布式的、非模块化的句法和语义底层。
The activations of language transformers like GPT-2 have been shown to linearly map onto brain activity during speech comprehension. However, the nature of these activations remains largely unknown and presumably conflate distinct linguistic classes. Here, we propose a taxonomy to factorize the high-dimensional activations of language models into four combinatorial classes: lexical, compositional, syntactic, and semantic representations. We then introduce a statistical method to decompose, through the lens of GPT-2's activations, the brain activity of 345 subjects recorded with functional magnetic resonance imaging (fMRI) during the listening of ~4.6 hours of narrated text. The results highlight two findings. First, compositional representations recruit a more widespread cortical network than lexical ones, and encompass the bilateral temporal, parietal and prefrontal cortices. Second, contrary to previous claims, syntax and semantics are not associated with separated modules, but, instead, appear to share a common and distributed neural substrate. Overall, this study introduces a versatile framework to isolate, in the brain activity, the distributed representations of linguistic constructs.
研究动机与目标
- 提出一种清晰且可操作的分类法,用以区分人工与生物神经网络中的词汇、组合、句法和语义表示。
- 开发一种从深度语言模型中提取句法表示并相应地分解脑表示的方法。
- 评估共享的脑-模型表示是局部化的还是在句法与语义层面分布的。
- 评估在自然叙事聆听期间,GPT-2 激活(按语言类别分解)与大规模队列的 fMRI 信号之间的映射。
提出的方法
- 定义一个五点分类法,在分布式激活中区分词汇、组合、句法和语义表示。
- 通过合成具有相同句法结构的句子并对 GPT-2 激活进行平均(overline{X}),来分离句法表示。
- 使用带岭回归和 FIR 延迟的时空编码模型,将模型激活 X(以及句法提取的 overline{X})映射到脑信号 Y。
- 将脑分数分解为跨 GPT-2 层的词汇、组合、句法和语义分量,并与句法合成基线进行比较。
- 将该方法应用于 Narratives fMRI 数据集(345 名受试者,约 4 小时的故事),并在不同层和架构上进行泛化测试。
实验结果
研究问题
- RQ1一个稳健的分类法是否能够在 GPT-2 与大脑活动中解开词汇、组合、句法和语义表示的结构?
- RQ2句法和语义表示是映射到独立的局部脑模块,还是映射到分布式底物?
- RQ3组合表示在脑映射上与词汇表示如何比较,它们分布在哪些区域?
- RQ4脑-模型映射是否在不同层之间以及在不同的变换器架构之间保持稳定?
主要发现
- 组合性表示比词汇表示动员了更广泛的皮质网络,涉及双侧颞叶、顶叶及前额叶皮质。
- 句法与语义并非局部分离的模块;它们共享一个共同且分布式的神经底物。
- 上下文(深层)层(例如 GPT-2 第 9 层)比词汇嵌入产生更高的脑分数,反映对语言结构的更强编码。
- 通过对句法匹配的合成句子进行激活平均,可以从 GPT-2 中提取句法表示,并且对脑映射仍具有信息性。
- 组合句法与组合语义在大脑中的参与是分布式的, peak 于包括颞叶和前额叶皮质、扣带皮层、上顶叶缘和中前额叶区等区域。
- 研究结果在不同层和架构上具有一般化性,中等 GPT-2 层往往产生最强的脑分数,并在其他变换器模型中观察到相似趋势。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。