東京工業大学 · 情報科学
東京工業大学のベラール教授研究室は、バイオインフォマティクスとデータサイエンスの分野において、がんの遺伝子発現データを活用した分類・予測モデルの構築を主眼としています。特に、小標本・高次元なマイクロアレイデータにおける分類器の評価手法や、AUCの解釈、交差検証におけるバイアスの是正、生存予後との統合的解析(生存木)に注力しています。また、仮説検定の限界を乗り越えるための効果量の信頼区間推定など、統計的妥当性を重視した機械学習の評価フレームワークの構築も進めています。
Figures are computed from collected data and may differ slightly.
The receiver operating characteristic (ROC) has emerged as the gold standard for assessing and comparing the performance of classifiers in a wide range of disciplines including the life sciences. ROC curves are frequently summarized in a single scalar, the area under the curve (AUC). This article discusses the caveats and pitfalls of ROC analysis in clinical microarray research, particularly in relation to (i) the interpretation of AUC (especially a value close to 0.5); (ii) model comparisons ba
Gene expression profiling by microarray technology has been successfully applied to classification and diagnostic prediction of cancers. Various machine learning and data mining methods are currently used for classifying gene expression data. However, these methods have not been developed to address the specific requirements of gene microarray analysis. First, microarray data is characterized by a high-dimensional feature space often exceeding the sample space dimensionality by a factor of 100 o
Commonly used accuracy-based performance values, with or without confidence intervals, are inadequate for comparing classifiers for small-sample data. We present a statistical methodology that avoids bias in cross-validated model selection in the context of small-sample scenarios. This methodology is valid for both k-fold cross-validation and repeated random sampling.
Deep learning in bioinformatics and biomedicine Daniel Berrar, Daniel Berrar Data Science Laboratory, Department of Information and Communications Engineering, Tokyo Institute of Technology, Ookayama, Tokyo 152-8550, Japan Corresponding author: Daniel Berrar, Data Science Laboratory, Department of Information and Communications Engineering, Tokyo Institute of Technology, Ookayama, Tokyo 152-8550, Japan. E-mail: daniel.berrar@ict.e.titech.ac.jp Search for other works by this author on: Oxford Aca
We present survival trees as an exploratory tool for revealing new insights into gene expression profiles in combination with clinical patient data. Survival trees partition the patient data studied into groups with similar survival outcomes and identify characteristic genetic profiles within these groups. We demonstrate the application of survival trees in a study involving the expression profiles of 3,588 genes in 211 lung adenocarcinoma patients. The survival tree identified a group of early-
Null hypothesis significance tests and their p-values currently dominate the statistical evaluation of classifiers in machine learning. Here, we discuss fundamental problems of this research practice. We focus on the problem of comparing multiple fully specified classifiers on a small-sample test set. On the basis of the method by Quesenberry and Hurst, we derive confidence intervals for the effect size, i.e. the difference in true classification performance. These confidence intervals disentang
Open papers in the app to read, cite, and organize with AI.