[论文解读] Autoreject: Automated artifact rejection for MEG and EEG data
Autoreject 是一种数据驱动的自动化方法,通过交叉验证为每个受试者和传感器优化峰峰值阈值,实现对 MEG 和 EEG 数据中坏试次和坏传感器的检测与修复。该方法通过用稳健、基于物理原理且完全自动化的流程替代人工伪影剔除,提升了结果的可重复性和可扩展性,同时保留数据并减少人为偏差。
We present an automated algorithm for unified rejection and repair of bad trials in magnetoencephalography (MEG) and electroencephalography (EEG) signals. Our method capitalizes on cross-validation in conjunction with a robust evaluation metric to estimate the optimal peak-to-peak threshold -- a quantity commonly used for identifying bad trials in M/EEG. This approach is then extended to a more sophisticated algorithm which estimates this threshold for each sensor yielding trial-wise bad sensors. Depending on the number of bad sensors, the trial is then repaired by interpolation or by excluding it from subsequent analysis. All steps of the algorithm are fully automated thus lending itself to the name Autoreject. In order to assess the practical significance of the algorithm, we conducted extensive validation and comparison with state-of-the-art methods on four public datasets containing MEG and EEG recordings from more than 200 subjects. Comparison include purely qualitative efforts as well as quantitatively benchmarking against human supervised and semi-automated preprocessing pipelines. The algorithm allowed us to automate the preprocessing of MEG data from the Human Connectome Project (HCP) going up to the computation of the evoked responses. The automated nature of our method minimizes the burden of human inspection, hence supporting scalability and reliability demanded by data analysis in modern neuroscience.
研究动机与目标
- 解决由于伪影剔除导致的预处理中缺乏共识且人工负担高的问题。
- 开发一种自动化、数据驱动的方法,替代主观性强、依赖专家判断的试次和传感器剔除。
- 通过插值修复坏试次而非直接剔除,最大限度减少数据损失,保留昂贵获取的数据。
- 确保与现有分析流程兼容,包括噪声归一化和脑源定位,无需修改算法。
- 为大规模数据集(如人类连接组计划的数据)提供可扩展、可重复且无偏差的预处理。
提出的方法
- 使用交叉验证和稳健的评估指标(训练集均值与验证集中位数之间的均方根误差)估计每个受试者的最优全局峰峰值阈值。
- 将全局阈值方法扩展至传感器特定的阈值,实现对每个试次中坏传感器的检测。
- 采用两阶段修复策略:若受影响的坏传感器数量少于阈值,则插值修复;否则剔除整个试次。
- 通过数据驱动的交叉验证框架校准所有参数,该框架对异常值具有鲁棒性并避免过拟合。
- 与标准 M/EEG 流程无缝集成,包括需要噪声归一化协方差估计的流程,无需修改下游处理步骤。
- 采用贝叶斯超参数优化方法,调节触发试次剔除的坏传感器数量,以在数据保留与噪声抑制之间实现最佳平衡。
实验结果
研究问题
- RQ1数据驱动的自动化方法能否在 MEG 和 EEG 数据中实现与人工标注伪影剔除相当或更优的性能?
- RQ2在检测坏试次的敏感性和特异性方面,传感器特定阈值与全局阈值相比表现如何?
- RQ3Autoreject 在多大程度上能保持数据质量,同时最大限度减少大规模 MEG/EEG 研究中对专家干预的依赖?
- RQ4该算法是否可无缝集成到标准 M/EEG 流程中,而无需修改下游步骤(如脑源定位或噪声归一化)?
- RQ5Autoreject 在具有不同采集设置和伪影特征的异质性数据集中的表现如何?
主要发现
- 在涵盖超过 200 名受试者的四个公开数据集上,Autoreject 的性能达到或优于现有最先进方法。
- 该算法成功自动化处理了人类连接组计划(HCP)的 MEG 数据,实现了无需人工检查的端到端诱发反应计算。
- 传感器特定阈值显著提升了检测准确性,这通过不同受试者和数据集中最优阈值的经验变异性得到验证。
- 通过修复数据而非丢弃试次,Autoreject 在某些数据集中将数据损失减少了高达 30%,相比传统剔除策略。
- 该方法对异常值和采集设置的差异表现出鲁棒性,适用于多中心、多扫描仪数据的整合。
- 通过交叉验证校准参数,Autoreject 减少了人为偏差,提升了可重复性,其结果在不同数据集和分析流程中均具有一致的可复现性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。