Skip to main content
QUICK REVIEW

[论文解读] How does longer temporal context enhance multimodal narrative video processing in the brain?

Prachi Jindal, Anant Khandelwal|arXiv (Cornell University)|Feb 7, 2026
Action Observation and Synchronization被引用 0
一句话总结

这项研究表明更长的时间上下文(3–12 秒片段)在自然电影观看中提升了多模态视频–音频大语言模型的脑对齐,且呈现ROI和层级相关模式,而单模态视频模型几乎没有提升。

ABSTRACT

Understanding how humans and artificial intelligence systems process complex narrative videos is a fundamental challenge at the intersection of neuroscience and machine learning. This study investigates how the temporal context length of video clips (3--12 s clips) and the narrative-task prompting shape brain-model alignment during naturalistic movie watching. Using fMRI recordings from participants viewing full-length movies, we examine how brain regions sensitive to narrative context dynamically represent information over varying timescales and how these neural patterns align with model-derived features. We find that increasing clip duration substantially improves brain alignment for multimodal large language models (MLLMs), whereas unimodal video models show little to no gain. Further, shorter temporal windows align with perceptual and early language regions, while longer windows preferentially align higher-order integrative regions, mirrored by a layer-to-cortex hierarchy in MLLMs. Finally, narrative-task prompts (multi-scene summary, narrative summary, character motivation, and event boundary detection) elicit task-specific, region-dependent brain alignment patterns and context-dependent shifts in clip-level tuning in higher-order regions. Together, our results position long-form narrative movies as a principled testbed for probing biologically relevant temporal integration and interpretable representations in long-context MLLMs.

研究动机与目标

  • 理解人类与AI处理长时叙事视频的方式及时间上下文在脑–模型对齐中的作用
  • 在不同片段长度下评估多模态视频–音频LLMs与单模态视频模型的脑预测性
  • 探讨叙事任务提示如何塑造区域特异的脑对齐与模型层之间的对应关系
  • 识别哪些视频片段与叙事提示最能驱动体素反应以理解情境相关的表征

提出的方法

  • 使用两种预训练的视频–音频MLLM(Qwen-2.5-Omni 与 DATE)和两种单模态基线模型(TimeSFormer、VideoMAE),在滑动时间窗口(3、6、9、12 s)下以1.49 s步幅生成表示
  • 从所有Transformer层提取表示,并对每个窗口和任务指令的标记进行平均
  • 构建体素级编码模型(自助重复抽样岭回归)以从刺激表示预测fMRI反应
  • 估计跨被试预测准确性以实现不同被试间的脑对齐归一化
  • 评估四种叙事任务(人物动机、事件边界检测、多场景摘要、叙事摘要)作为提示以获得任务特异表示
  • 分析层级与ROI特异的对齐以检验时间梯度和皮质层级结构
Figure 1: Leveraging temporal video context of different durations ( $X_{\text{windows}}$ ) with unimodal and multimodal models for brain encoding with a diverse set of instructions (prompts). We experiment with 4 narrative video understanding tasks: character motivation, event boundary detection, m
Figure 1: Leveraging temporal video context of different durations ( $X_{\text{windows}}$ ) with unimodal and multimodal models for brain encoding with a diverse set of instructions (prompts). We experiment with 4 narrative video understanding tasks: character motivation, event boundary detection, m

实验结果

研究问题

  • RQ1RQ1 在自然电影观看中,增加时间上下文长度如何影响多模态 versus 单模态视频模型的脑预测性?
  • RQ2RQ2 哪些脑区表现出对更长上下文的收益或最佳上下文长度的转变,以及这些与MLLM层表示有何关系?
  • RQ3RQ3 叙事任务提示如何影响脑对齐,是否在ROI层面呈现出区分?
  • RQ4RQ4 哪些视频片段在不同情境和任务下最强地驱动体素反应,以及不同ROI的模式如何变化?

主要发现

  • 更长的时间上下文显著提高了视频–音频MLLMs的脑对齐(Qwen-2.5-Omni 约26%相对提升,DATE约19%),而对单模态基线几乎无变化
  • 较长的窗口(12 s)更偏向对高级语义区域的对齐(如PCC、dmPFC),而中等长度的窗口(6 s)则偏向感知与早期语言区域(如PTL)
  • 叙事任务提示产生任务特异且与ROI相关的对齐模式;叙事与多场景摘要激活较高级别区域,人物动机激活时序语言区域,事件边界检测更局部化
  • 层级分析显示皮质语言层级结构:更深的层与更高级脑区域对齐,而更早的层与感官区域在不同时间上下文中保持一致
  • 视觉ROIs在各上下文中表现出稳定的视频剪辑偏好,但更高层区域(如AG、PCC)会随上下文和提示而转变
  • 在体素响应上,最激活的片段在视觉区域保持稳定,但在更高层区域会随上下文增长和提示而转变,显示语义敏感性随情境而变化
Figure 2: Average normalized brain alignment as a function of temporal window length (3 to 12s) for MLLMs, and unimodal video baselines. MLLMs show increasing alignment with longer windows, while unimodal video models remain approximately constant. Error bars denote variability across subjects (mean
Figure 2: Average normalized brain alignment as a function of temporal window length (3 to 12s) for MLLMs, and unimodal video baselines. MLLMs show increasing alignment with longer windows, while unimodal video models remain approximately constant. Error bars denote variability across subjects (mean

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。