Skip to main content
QUICK REVIEW

[Paper Review] How does longer temporal context enhance multimodal narrative video processing in the brain?

Prachi Jindal, Anant Khandelwal|arXiv (Cornell University)|Feb 7, 2026
Action Observation and Synchronization0 citations
TL;DR

This study shows that longer temporal context (3–12 s clips) improves brain alignment for multimodal video–audio LLMs during naturalistic movie viewing, with ROI- and layer-dependent patterns, while unimodal video models show little gain.

ABSTRACT

Understanding how humans and artificial intelligence systems process complex narrative videos is a fundamental challenge at the intersection of neuroscience and machine learning. This study investigates how the temporal context length of video clips (3--12 s clips) and the narrative-task prompting shape brain-model alignment during naturalistic movie watching. Using fMRI recordings from participants viewing full-length movies, we examine how brain regions sensitive to narrative context dynamically represent information over varying timescales and how these neural patterns align with model-derived features. We find that increasing clip duration substantially improves brain alignment for multimodal large language models (MLLMs), whereas unimodal video models show little to no gain. Further, shorter temporal windows align with perceptual and early language regions, while longer windows preferentially align higher-order integrative regions, mirrored by a layer-to-cortex hierarchy in MLLMs. Finally, narrative-task prompts (multi-scene summary, narrative summary, character motivation, and event boundary detection) elicit task-specific, region-dependent brain alignment patterns and context-dependent shifts in clip-level tuning in higher-order regions. Together, our results position long-form narrative movies as a principled testbed for probing biologically relevant temporal integration and interpretable representations in long-context MLLMs.

Motivation & Objective

  • Motivate understanding of how humans and AI process long-form narrative videos and the role of temporal context in brain–model alignment.
  • Assess brain predictivity of multimodal video–audio LLMs versus unimodal video models across varying clip lengths.
  • Investigate how narrative-task prompts shape region-specific brain alignment and model-layer correspondences.
  • Identify which video clips and prompts most strongly drive voxel responses to understand context-dependent representations.

Proposed method

  • Use two pretrained video–audio MLLMs (Qwen-2.5-Omni and DATE) and two unimodal baselines (TimeSFormer, VideoMAE) to generate representations from sliding temporal windows (3, 6, 9, 12 s) with 1.49 s stride.
  • Extract representations from all Transformer layers and average across tokens per window and task instruction.
  • Build voxel-wise encoding models (bootstrap ridge regression) to predict fMRI responses from stimulus representations.
  • Estimate cross-subject prediction accuracy to normalize brain alignment across subjects.
  • Evaluate four narrative tasks (Character Motivation, Event Boundary Detection, Multi-Scene Summary, Narrative Summary) as prompts to obtain task-specific representations.
  • Analyze layer-wise and ROI-specific alignment to examine temporal gradients and cortical hierarchies.
Figure 1: Leveraging temporal video context of different durations ( $X_{\text{windows}}$ ) with unimodal and multimodal models for brain encoding with a diverse set of instructions (prompts). We experiment with 4 narrative video understanding tasks: character motivation, event boundary detection, m
Figure 1: Leveraging temporal video context of different durations ( $X_{\text{windows}}$ ) with unimodal and multimodal models for brain encoding with a diverse set of instructions (prompts). We experiment with 4 narrative video understanding tasks: character motivation, event boundary detection, m

Experimental results

Research questions

  • RQ1RQ1 How does increasing temporal context length affect brain predictivity for multimodal versus unimodal video models during naturalistic movie watching?
  • RQ2RQ2 Which brain regions show gains or shifts in optimal context length, and how do these relate to MLLM layer representations?
  • RQ3RQ3 How do narrative-task prompts influence brain alignment, and do they dissociate into ROI-specific patterns?
  • RQ4RQ4 Which video clips most strongly drive voxel responses across contexts and tasks, and how do patterns vary by ROI?

Key findings

  • Longer temporal context improves brain alignment for video–audio MLLMs (≈26% relative gain for Qwen-2.5-Omni and ≈19% for DATE) but not for unimodal baselines (negligible change).
  • Long windows (12 s) preferentially align higher-order semantic regions (e.g., PCC, dmPFC) while intermediate windows (6 s) favor perceptual and early language regions (e.g., PTL).
  • Narrative-task prompts yield task-specific, ROI-dependent alignment patterns; Narrative and Multi-scene Summaries engage higher-order regions, Character Motivation engages temporal language regions, and Event Boundary Detection is more localized.
  • Layer-wise analysis shows a cortical language hierarchy: deeper layers align with higher-order brain regions, while earlier layers align with sensory regions, across temporal contexts.
  • Visual ROIs show stable clip preferences across contexts, whereas higher-order regions (AG, PCC) shift with context and prompts.
  • Top-activating clips for voxel responses are stable in visual regions but shift in higher-order regions as context grows, indicating context-dependent semantic sensitivity.
Figure 2: Average normalized brain alignment as a function of temporal window length (3 to 12s) for MLLMs, and unimodal video baselines. MLLMs show increasing alignment with longer windows, while unimodal video models remain approximately constant. Error bars denote variability across subjects (mean
Figure 2: Average normalized brain alignment as a function of temporal window length (3 to 12s) for MLLMs, and unimodal video baselines. MLLMs show increasing alignment with longer windows, while unimodal video models remain approximately constant. Error bars denote variability across subjects (mean

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.