Tokyo Institute of Technology · Computer Science
Professor Daniel Berrar's research lab specializes in data science applications within bioinformatics and biomedicine, with a strong focus on machine learning, statistical modeling, and high-dimensional data analysis. The lab investigates challenges in small-sample and high-dimensional settings—common in genomics and clinical microarray data—emphasizing robust performance evaluation, model selection, and reliable inference. Key research directions include the development of statistical methodologies for classifier evaluation (e.g., AUC analysis, confidence intervals for effect sizes), survival analysis using gene expression data, and the application of deep learning in biomedical contexts. The lab also addresses methodological pitfalls in classical hypothesis testing and promotes more informative alternatives to p-values in machine learning research.
Figures are computed from collected data and may differ slightly.
The receiver operating characteristic (ROC) has emerged as the gold standard for assessing and comparing the performance of classifiers in a wide range of disciplines including the life sciences. ROC curves are frequently summarized in a single scalar, the area under the curve (AUC). This article discusses the caveats and pitfalls of ROC analysis in clinical microarray research, particularly in relation to (i) the interpretation of AUC (especially a value close to 0.5); (ii) model comparisons ba
Gene expression profiling by microarray technology has been successfully applied to classification and diagnostic prediction of cancers. Various machine learning and data mining methods are currently used for classifying gene expression data. However, these methods have not been developed to address the specific requirements of gene microarray analysis. First, microarray data is characterized by a high-dimensional feature space often exceeding the sample space dimensionality by a factor of 100 o
Commonly used accuracy-based performance values, with or without confidence intervals, are inadequate for comparing classifiers for small-sample data. We present a statistical methodology that avoids bias in cross-validated model selection in the context of small-sample scenarios. This methodology is valid for both k-fold cross-validation and repeated random sampling.
Deep learning in bioinformatics and biomedicine Daniel Berrar, Daniel Berrar Data Science Laboratory, Department of Information and Communications Engineering, Tokyo Institute of Technology, Ookayama, Tokyo 152-8550, Japan Corresponding author: Daniel Berrar, Data Science Laboratory, Department of Information and Communications Engineering, Tokyo Institute of Technology, Ookayama, Tokyo 152-8550, Japan. E-mail: daniel.berrar@ict.e.titech.ac.jp Search for other works by this author on: Oxford Aca
We present survival trees as an exploratory tool for revealing new insights into gene expression profiles in combination with clinical patient data. Survival trees partition the patient data studied into groups with similar survival outcomes and identify characteristic genetic profiles within these groups. We demonstrate the application of survival trees in a study involving the expression profiles of 3,588 genes in 211 lung adenocarcinoma patients. The survival tree identified a group of early-
Null hypothesis significance tests and their p-values currently dominate the statistical evaluation of classifiers in machine learning. Here, we discuss fundamental problems of this research practice. We focus on the problem of comparing multiple fully specified classifiers on a small-sample test set. On the basis of the method by Quesenberry and Hurst, we derive confidence intervals for the effect size, i.e. the difference in true classification performance. These confidence intervals disentang
Open papers in the app to read, cite, and organize with AI.