Skip to main content
QUICK REVIEW

[论文解读] Inducing brain-relevant bias in natural language processing models

Dan Schwartz, Mariya Toneva|arXiv (Cornell University)|Oct 29, 2019
Neurobiology of Language and Bilingualism参考文献 27被引用 22
一句话总结

本文提出通过使用脑活动记录(fMRI 和 MEG)对 BERT 进行微调,以生成更能捕捉与大脑相关语言处理的表征。微调后,该模型在跨参与者预测 fMRI 活动方面表现更优,且优于仅使用 fMRI 的模型,表明基于大脑的表征学习可提升神经活动预测性能,同时不损害自然语言处理(NLP)性能。

ABSTRACT

Progress in natural language processing (NLP) models that estimate representations of word sequences has recently been leveraged to improve the understanding of language processing in the brain. However, these models have not been specifically designed to capture the way the brain represents language meaning. We hypothesize that fine-tuning these models to predict recordings of brain activity of people reading text will lead to representations that encode more brain-activity-relevant language information. We demonstrate that a version of BERT, a recently introduced and powerful language model, can improve the prediction of brain activity after fine-tuning. We show that the relationship between language and brain activity learned by BERT during this fine-tuning transfers across multiple participants. We also show that, for some participants, the fine-tuned representations learned from both magnetoencephalography (MEG) and functional magnetic resonance imaging (fMRI) are better for predicting fMRI than the representations learned from fMRI alone, indicating that the learned representations capture brain-activity-relevant information that is not simply an artifact of the modality. While changes to language representations help the model predict brain activity, they also do not harm the model's ability to perform downstream NLP tasks. Our findings are notable for research on language understanding in the brain.

研究动机与目标

  • 探究是否可通过在脑活动记录上微调类似 BERT 的预训练语言模型,生成更能预测语言相关神经反应的表征。
  • 确定通过多个参与者和多种神经影像模态(fMRI 和 MEG)学习的表征是否能在个体间和不同记录类型间实现泛化。
  • 评估基于大脑的微调是否能提升 fMRI 活动预测性能,超越仅使用 fMRI 数据所能达到的效果。

提出的方法

  • 使用线性头对 BERT 进行微调,以从对应文本输入中预测 fMRI 和 MEG 记录,其中 MEG 预测基于内容词,fMRI 预测基于 [CLS] 标记嵌入。
  • 采用多任务学习,在微调过程中同时优化 fMRI 和 MEG 预测,以利用两种模态的互补优势。
  • 在多个参与者的数据上进行训练,以促进脑相关表征在个体间的泛化能力。
  • 使用可解释方差比例(PoV)评估模型性能,比较微调模型与原始 BERT 的表现。
  • 分析变化最显著的样本中特征分布的变化,以识别在微调过程中驱动表征变化的语言特征。
  • 使用自举重采样法估计高变化与低变化样本中特征流行度比较的标准误。

实验结果

研究问题

  • RQ1在脑活动记录上微调 BERT 是否能提升其对自然语言刺激 fMRI 反应的预测能力?
  • RQ2基于大脑的表征学习是否能在不同参与者和神经影像模态之间实现迁移?
  • RQ3在 MEG 数据上微调的表征是否能提升 fMRI 预测性能,超越仅在 fMRI 上微调的模型?
  • RQ4在基于大脑的微调过程中,BERT 表征的变化是否与特定语言特征(如运动或情绪标签)相对应?
  • RQ5基于大脑的微调带来的性能提升是否局限于特定脑区,还是广泛分布于语言相关区域?

主要发现

  • 微调后的 BERT 模型显著提升了跨参与者对 fMRI 活动的预测能力,其中 MEG 迁移模型在大多数参与者的语言相关区域中表现优于仅使用 fMRI 微调的模型。
  • 微调后的表征在参与者间具有泛化能力,表明模型学习到了共享的、与大脑相关的语言模式。
  • 对于部分参与者,同时在 MEG 和 fMRI 数据上微调的表征比仅在 fMRI 上微调的表征更准确地预测 fMRI 活动,表明 MEG 捕捉到了非冗余的、与大脑相关的有效信息。
  • 微调过程中变化最显著的样本显示出更高比例的与运动相关的标签(如“move”)和祈使句语言,表明这些语言特征在表征适应中起核心作用。
  • 在基于大脑的微调后,模型在下游 NLP 任务中仍保持强劲性能,证实大脑相关偏差并未损害其核心语言理解能力。
  • 特征分布分析显示,动词和限定词是受微调影响最显著的语言类别之一,表明它们在神经语言处理中具有重要意义。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。