Skip to main content
QUICK REVIEW

[论文解读] A Novel Approach for Stable Selection of Informative Redundant Features from High Dimensional fMRI Data

Yilun Wang, Zhiqiang Li|arXiv (Cornell University)|Jun 27, 2015
Face and Expression Recognition参考文献 58被引用 7
一句话总结

本文提出了一种新颖的特征选择方法,结合稳定性选择与弹性网络正则化,以提升高维fMRI数据中生物标志物发现的稳定性、鲁棒性和可解释性。通过选择以往在传统方法中被忽视的信息冗余特征,该方法在噪声标签和数据变异条件下,实现了对假阳性与漏检的优越控制,其有效性已在合成数据及多中心ADHD fMRI数据集上得到验证。

ABSTRACT

Feature selection is among the most important components because it not only helps enhance the classification accuracy, but also or even more important provides potential biomarker discovery. However, traditional multivariate methods is likely to obtain unstable and unreliable results in case of an extremely high dimensional feature space and very limited training samples, where the features are often correlated or redundant. In order to improve the stability, generalization and interpretations of the discovered potential biomarker and enhance the robustness of the resultant classifier, the redundant but informative features need to be also selected. Therefore we introduced a novel feature selection method which combines a recent implementation of the stability selection approach and the elastic net approach. The advantage in terms of better control of false discoveries and missed discoveries of our approach, and the resulted better interpretability of the obtained potential biomarker is verified in both synthetic and real fMRI experiments. In addition, we are among the first to demonstrate the robustness of feature selection benefiting from the incorporation of stability selection and also among the first to demonstrate the possible unrobustness of the classical univariate two-sample t-test method. Specifically, we show the robustness of our feature selection results in existence of noisy (wrong) training labels, as well as the robustness of the resulted classifier based on our feature selection results in the existence of data variation, demonstrated by a multi-center attention-deficit/hyperactivity disorder (ADHD) fMRI data.

研究动机与目标

  • 解决传统多变量特征选择在高维fMRI数据中样本有限时的不稳定性与不可靠性问题。
  • 在训练标签存在噪声及多中心数据变异条件下,提升特征选择的鲁棒性。
  • 实现对信息冗余特征(此前被忽略)的选择,从而增强生物标志物的可解释性与分类器的泛化能力。
  • 展示经典单变量t检验在特征选择中的局限性,尤其是在标签噪声存在时的表现。
  • 提供一种稳定、可复现的方法,用于在复杂相关数据中识别潜在的神经影像生物标志物。

提出的方法

  • 将稳定性选择与弹性网络正则化相结合,联合控制高维特征空间中的假阳性与假阴性。
  • 通过子采样应用稳定性选择,以估计在多个随机训练数据子集上的特征选择稳定性。
  • 利用弹性网络的混合L1/L2正则化处理相关特征,并选择信息冗余特征。
  • 结合稳定性选择的选取频率与弹性网络的系数收缩,识别稳定且信息丰富的特征。
  • 利用稳定性选择的理论保证优化选择阈值,以控制错误率。
  • 在所选特征上训练最终分类器,以评估在数据变异与标签噪声下的泛化性能。

实验结果

研究问题

  • RQ1稳定性选择与弹性网络的混合方法是否能提升高维fMRI数据中特征选择的稳定性与可靠性?
  • RQ2引入信息冗余特征在多大程度上影响所识别生物标志物的可解释性与鲁棒性?
  • RQ3在标签噪声条件下,该方法相较于经典单变量t检验的性能优势有多大?
  • RQ4该方法在多中心fMRI数据存在变异时,其特征选择与最终分类器的鲁棒性如何?
  • RQ5该方法在真实神经影像应用中能否有效控制假阳性发现与漏检?

主要发现

  • 与传统多变量及单变量方法相比,所提方法显著减少了假阳性与漏检数量。
  • 该方法成功识别出信息冗余特征,增强了所获生物标志物集合的可解释性。
  • 即使训练标签中有高达20%被噪声污染,特征选择结果仍保持稳定与可靠。
  • 基于该方法所选特征训练的分类器,在来自不同中心的多个ADHD fMRI数据集中均表现出良好的泛化能力。
  • 本研究首次表明,经典单变量t检验在标签噪声下极为不鲁棒,而所提方法则保持了稳定的性能表现。
  • 将稳定性选择与弹性网络结合,显著提升了高维fMRI数据中特征选择流程的鲁棒性与可复现性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。