[Paper Review] Can Vision Replace Text in Working Memory? Evidence from Spatial n-Back in Vision-Language Models
The paper compares spatial n-back performance in text-based LLMs and vision-language models using matched text-grid and image-grid stimuli, finding a modality-dependent performance gap and evidence of recency-based strategies and interference in vision formats.
Working memory is a central component of intelligent behavior, providing a dynamic workspace for maintaining and updating task-relevant information. Recent work has used n-back tasks to probe working-memory-like behavior in large language models, but it is unclear whether the same probe elicits comparable computations when information is carried in a visual rather than textual code in vision-language models. We evaluate Qwen2.5 and Qwen2.5-VL on a controlled spatial n-back task presented as matched text-rendered or image-rendered grids. Across conditions, models show reliably higher accuracy and d' with text than with vision. To interpret these differences at the process level, we use trial-wise log-probability evidence and find that nominal 2/3-back often fails to reflect the instructed lag and instead aligns with a recency-locked comparison. We further show that grid size alters recent-repeat structure in the stimulus stream, thereby changing interference and error patterns. These results motivate computation-sensitive interpretations of multimodal working memory.
Motivation & Objective
- Motivate whether changing representational code (text vs. vision) preserves the intended working-memory computations in multimodal models.
- Assess how modality affects updating, temporal-binding, and interference control in a spatial n-back task.
- Provide process-level diagnostics to interpret WM-like behavior beyond end-point accuracy.
Proposed method
- Evaluate Qwen2.5 (text) and Qwen2.5-VL (vision-language) on a spatial n-back task with matched text-rendered and image-rendered grids.
- Use deterministic decoding to ensure trial-by-trial reproducibility.
- Compute standard WM metrics (accuracy, hit rate, false alarm rate, d′) and a trial-wise log-probability-based evidence score (s_t).
- Perform a lag-scan analysis by re-labeling trials with different lag definitions (k ∈ {1,2,3}) to assess whether evidence aligns with the instructed lag or a recency-based strategy.
- Manipulate grid size N (3,4,5,7) to study interference from recent repeats (lures) and its impact on discriminability.
- Assess robustness across model families (Llama3 and Qwen variants) and scales using text-grid and vision-grid inputs.
Experimental results
Research questions
- RQ1Does replacing text with vision in matched spatial n-back tasks yield the same WM-like computations or a modality-driven shift in strategy?
- RQ2How do temporal-context binding and interference control differ between text-grid and vision-grid representations in multimodal models?
- RQ3How do grid size and nominal load affect discriminability and interference patterns across modalities?
- RQ4Is performance primarily limited by a conservative decision criterion or by weak evidence separability under different lags?
- RQ5Do lag-based diagnostics reveal recency-locked processing rather than instructed lag matching in higher loads?
Key findings
- Across loads and grid sizes, text-grid input yields higher discriminability than text-grid vision, with vision-grid performing the worst.
- Sensitivity d′ decreases substantially with nominal load (1-back to 2/3-back) across conditions, and the vision-grid condition remains notably below text-grid baselines.
- Lag-scan analyses show that graded evidence often aligns with a 1-back recency-locked comparison rather than the instructed n-back lag, especially for n=2,3.
- Grid-size effects reveal that larger grids improve discriminability, while the vision-grid condition remains far below text-grid baselines; interference from recent repeats (lures) underpins much of the grid-size advantage.
- Process diagnostics show that vision-grid often yields a conservative bias (low hit rates, low false alarms) and weak evidence separability, with AUC peaking at k=1 rather than k=n.
- Robustness checks across LLM/VLM families (Llama3 and Qwen variants) show the same qualitative patterns: text-grid > vision-grid with modality order, and larger grids generally increasing d′, though magnitudes vary by model.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.