[Paper Review] CrossPT-EEG: A Benchmark for Cross-Participant and Cross-Time Generalization of EEG-based Visual Decoding
This paper introduces EEG-ImageNet, a large-scale EEG dataset with 16 participants and 4,000 image stimuli from ImageNet-21k, supporting multi-granularity (coarse and fine) labels. It establishes benchmarks for cross-participant and cross-time generalization in EEG-based visual decoding, achieving 60.88% object classification accuracy and 64.67% two-way identification in image reconstruction using CLIP-based models.
Exploring brain activity in relation to visual perception provides insights into the biological representation of the world. While functional magnetic resonance imaging (fMRI) and magnetoencephalography (MEG) have enabled effective image classification and reconstruction, their high cost and bulk limit practical use. Electroencephalography (EEG), by contrast, offers low cost and excellent temporal resolution, but its potential has been limited by the scarcity of large, high-quality datasets and by block-design experiments that introduce temporal confounds. To fill this gap, we present CrossPT-EEG, a benchmark for cross-participant and cross-time generalization of visual decoding from EEG. We collected EEG data from 16 participants while they viewed 4,000 images sampled from ImageNet, with image stimuli annotated at multiple levels of granularity. Our design includes two stages separated in time to allow cross-time generalization and avoid block-design artifacts. We also introduce benchmarks tailored to non-block design classification, as well as pre-training experiments to assess cross-time and cross-participant generalization. These findings highlight the dataset's potential to enhance EEG-based visual brain-computer interfaces, deepen our understanding of visual perception in biological systems, and suggest promising applications for improving machine vision models.
Motivation & Objective
- To address the lack of large-scale, high-quality EEG datasets with multi-granularity labels for visual neuroscience research.
- To enable cross-participant and cross-time generalization in EEG-based visual decoding by providing a comprehensive benchmark.
- To support the development of robust brain-computer interfaces (BCIs) and improve understanding of human visual perception through EEG.
- To facilitate the application of deep learning models to EEG signals for object classification and image reconstruction tasks.
- To overcome limitations of existing EEG datasets, which are small in scale and lack fine-grained categorization.
Proposed method
- Collecting EEG recordings from 16 subjects exposed to 4,000 images from ImageNet-21k, with 50 images per of 80 categories.
- Designing the dataset to support both coarse-grained (40 categories) and fine-grained (40 categories) classification tasks.
- Aligning EEG signals with image captions generated by the BLIP model to enable image reconstruction pipelines.
- Training and evaluating multiple deep learning models—including AlexNet, Inception, and CLIP—for object classification and image reconstruction.
- Using two-way identification as an evaluation metric for image reconstruction, comparing generated images against original and distractor images.
- Implementing sequential data splitting and short segment lengths to mitigate temporal data leakage and improve cross-time generalization.
Experimental results
Research questions
- RQ1Can EEG-based visual decoding models generalize across different participants using a large-scale EEG dataset with multi-granularity labels?
- RQ2What is the upper limit of performance for EEG-based object classification and image reconstruction on a standardized benchmark?
- RQ3How does model architecture, particularly vision transformers like CLIP, affect performance in EEG-to-image reconstruction tasks?
- RQ4To what extent does multi-granularity labeling improve the interpretability and utility of EEG-based visual decoding models?
- RQ5How can domain adaptation and transfer learning be enhanced using this large-scale EEG dataset for cross-participant generalization?
Key findings
- The best-performing model achieved 60.88% accuracy in 80-class object classification, demonstrating strong cross-participant generalization on EEG-ImageNet.
- The CLIP-based model achieved the highest two-way identification rate of 64.67% in image reconstruction, outperforming other architectures.
- Image reconstruction results showed that category-level information was preserved, but low-level details like color, shape, and position were inaccurately restored.
- The use of BLIP-generated captions for EEG-image alignment limited the precision of reconstruction, suggesting a need for more accurate signal alignment methods.
- Performance varied across participants, indicating individual differences in neural responses and the need for personalized or domain-adapted models.
- The dataset enables meaningful evaluation of cross-time and cross-participant generalization, offering a foundation for future BCI and neuroimaging research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.