[Paper Review] Joint-Embedding Masked Autoencoder for Self-supervised Learning of Dynamic Functional Connectivity from the Human Brain
The paper introduces ST-JEMA, a JEPA-inspired masked autoencoder for dynamic graphs derived from fMRI, enabling self-supervised learning of high-level spatio-temporal representations to predict phenotypes and psychiatric diagnoses with limited labels.
Graph Neural Networks (GNNs) have shown promise in learning dynamic functional connectivity for distinguishing phenotypes from human brain networks. However, obtaining extensive labeled clinical data for training is often resource-intensive, making practical application difficult. Leveraging unlabeled data thus becomes crucial for representation learning in a label-scarce setting. Although generative self-supervised learning techniques, especially masked autoencoders, have shown promising results in representation learning in various domains, their application to dynamic graphs for dynamic functional connectivity remains underexplored, facing challenges in capturing high-level semantic representations. Here, we introduce the Spatio-Temporal Joint Embedding Masked Autoencoder (ST-JEMA), drawing inspiration from the Joint Embedding Predictive Architecture (JEPA) in computer vision. ST-JEMA employs a JEPA-inspired strategy for reconstructing dynamic graphs, which enables the learning of higher-level semantic representations considering temporal perspectives, addressing the challenges in fMRI data representation learning. Utilizing the large-scale UK Biobank dataset for self-supervised learning, ST-JEMA shows exceptional representation learning performance on dynamic functional connectivity demonstrating superiority over previous methods in predicting phenotypes and psychiatric diagnoses across eight benchmark fMRI datasets even with limited samples and effectiveness of temporal reconstruction on missing data scenarios. These findings highlight the potential of our approach as a robust representation learning method for leveraging label-scarce fMRI data.
Motivation & Objective
- Leverage large-scale unlabeled fMRI data to learn robust dynamic functional connectivity representations.
- Develop a JEPA-inspired masked autoencoder to capture higher-level spatial and temporal semantics in dynamic graphs.
- Demonstrate improved downstream performance on phenotype and psychiatric diagnosis tasks with limited labels.
Proposed method
- Introduce ST-JEMA with dual encoders for context and target components to avoid trivial representations.
- Apply a JEPA-based loss to reconstruct latent target representations rather than raw features across space and time.
- Use block masking on node features and adjacency matrices to learn from contextual graph information at each time step.
- Employ a global context node representation to manage multiple masks efficiently during training.
- Combine spatial and temporal reconstruction losses, with node and adjacency matrix predictions, to learn dynamic graph semantics.
- Fine-tune on eight benchmark rs-fMRI datasets to validate representation quality for downstream tasks.
Experimental results
Research questions
- RQ1Can JEPA-inspired masked autoencoding improve representation learning for dynamic functional connectivity from unlabeled fMRI data?
- RQ2Do spatio-temporal joint embeddings enhance downstream prediction of phenotypes and psychiatric diagnoses with limited labeled data?
- RQ3How does temporal reconstruction contribute to handling missing data scenarios in dynamic fMRI?
- RQ4What is the impact of large-scale unlabeled data (UK Biobank) on backbone GNN representations for fMRI?
- RQ5Does the proposed method outperform static-graph or contrastive SSL baselines on eight benchmarks?
Key findings
- ST-JEMA outperforms prior SSL methods in downstream tasks on eight benchmark fMRI datasets.
- Utilizing temporal dynamics improves node representations in the GNN encoder compared to baselines.
- Temporal reconstruction aids performance in data-scarce clinical settings, particularly for psychiatric diagnosis classification.
- Temporal reconstruction remains effective under missing data scenarios.
- Large-scale unlabeled data from UK Biobank supports stronger backbone representations for downstream tasks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.