Skip to main content
QUICK REVIEW

[論文レビュー] Gene selection for cancer classification using a hybrid of univariate and multivariate feature selection methods

Min Xu, Rudy Setiono|arXiv (Cornell University)|Jun 5, 2015
Gene expression and cancer classification参考文献 22被引用数 5
ひとこと要約

本稿では、がん分類の精度を向上させながら少ない遺伝子数で実現するため、単変量の最尤法(LIK)と多変量の再帰的特徴量削除法(RFE)を組み合わせたハイブリッド遺伝子選択手法を提案する。初期の遺伝子ランク付けにLIKを活用し、反復的な特徴量削除にRFEを適用することで、ノイズへの感受性と計算コストを低減しつつ、白血球およびSRBCTデータセットで優れたもしくは同等の精度を達成した。

ABSTRACT

Various approaches to gene selection for cancer classification based on microarray data can be found in the literature and they may be grouped into two categories: univariate methods and multivariate methods. Univariate methods look at each gene in the data in isolation from others. They measure the contribution of a particular gene to the classification without considering the presence of the other genes. In contrast, multivariate methods measure the relative contribution of a gene to the classification by taking the other genes in the data into consideration. Multivariate methods select fewer genes in general. However, the selection process of multivariate methods may be sensitive to the presence of irrelevant genes, noises in the expression and outliers in the training data. At the same time, the computational cost of multivariate methods is high. To overcome the disadvantages of the two types of approaches, we propose a hybrid method to obtain gene sets that are small and highly discriminative. We devise our hybrid method from the univariate Maximum Likelihood method (LIK) and the multivariate Recursive Feature Elimination method (RFE). We analyze the properties of these methods and systematically test the effectiveness of our proposed method on two cancer microarray datasets. Our experiments on a leukemia dataset and a small, round blue cell tumors dataset demonstrate the effectiveness of our hybrid method. It is able to discover sets consisting of fewer genes than those reported in the literature and at the same time achieve the same or better prediction accuracy.

研究の動機と目的

  • がん分類における単変量および多変量遺伝子選択手法の限界を解決すること。
  • 多変量手法における不要な遺伝子、ノイズ、外れ値への感受性を低減すること。
  • 多変量アプローチの高い計算コストを低減すること。
  • 分類精度を維持または向上させつつ、より小さな高 discriminative 遺伝子セットを特定すること。
  • 単変量および多変量手法を統合し、相乗効果を生むハイブリッドフレームワークを構築すること。

提案手法

  • 本手法は、個々の遺伝子が分類に与える寄与度に基づき、初期の遺伝子ランク付けに単変量の最尤法(LIK)を用いる。
  • 遺伝子間の相互作用を考慮しながら、反復的に関連性の低い遺伝子を削除する多変量の再帰的特徴量削減法(RFE)を適用する。
  • 遺伝子選択は、まずLIKによる遺伝子の順位付けを行い、その後RFEによる多変量的関連性に基づくセットの最適化を実施する。
  • ハイブリッド手法は、計算コストの高いRFE処理の前にLIKによる早期フィルタリングにより、ノイズや不要な特徴の影響を低減する。
  • 本手法は、急性白血球および小円形青色腫瘍(SRBCT)の2つのマイクロアレイデータセットを用いて評価された。
  • 特徴量選択は、主に分類精度を性能指標として検証された。

実験結果

リサーチクエスチョン

  • RQ1単変量および多変量手法を統合したハイブリッド手法は、がん分類の遺伝子選択を改善できるか?
  • RQ2LIKとRFEを統合することで、遺伝子数を削減しながら分類精度を維持または向上できるか?
  • RQ3提案手法は、ノイズおよび外れ値に対して従来の単体の多変量手法と比較して、どの程度高いロバスト性を示すか?
  • RQ4純粋な多変量手法と比較して、ハイブリッド手法は計算コストをどの程度低減できるか?
  • RQ5本手法は、既存の文献で報告されたものと比較して、より小さく、より特徴的な遺伝子セットを特定できるか?

主な発見

  • ハイブリッド手法は、白血球およびSRBCTデータセットの両方で、最先端の手法と同等または優れた分類精度を達成した。
  • 従来の報告結果と比較して、顕著に少ない遺伝子数を採用しながらも、高い予測性能を維持した。
  • 白血球データセットでは、過去の研究で報告されたより大きな遺伝子セットと同等または上回る精度を示す小さな遺伝子セットを同定した。
  • SRBCTデータセットでは、個別の単変量および多変量手法と比較して、遺伝子セットのサイズと精度の両面で本手法が優れた性能を示した。
  • LIKとRFEの統合により、単体の多変量手法と比較してノイズおよび外れ値へのロバスト性が向上した。
  • LIKによる事前のフィルタリングのおかげで、より高コストなRFEフェーズに渡される遺伝子数が制限され、計算コストが低減した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。