Tokyo Institute of Technology · 컴퓨터과학
다니엘 베라르 교수의 연구실은 생물정보학과 의료 데이터 과학 분야에서 주로 활동하며, 특히 유전자 발현 데이터와 임상 데이터를 융합한 분류 및 예측 모델 개발에 초점을 맞추고 있습니다. 고차원적이고 노이즈가 많은 마이크로어레이 및 유전자 발현 데이터를 대상으로 한 정밀한 분류 성능 평가, 특히 AUC 기반 평가와 교차검증에서의 편향 문제 해결에 대한 통계적 방법론을 개발하고 있습니다. 또한 생존 분석을 위한 생존 트리 모델과 딥러닝 기반 생물의학 응용에 대해서도 연구를 확장하고 있습니다.
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
The receiver operating characteristic (ROC) has emerged as the gold standard for assessing and comparing the performance of classifiers in a wide range of disciplines including the life sciences. ROC curves are frequently summarized in a single scalar, the area under the curve (AUC). This article discusses the caveats and pitfalls of ROC analysis in clinical microarray research, particularly in relation to (i) the interpretation of AUC (especially a value close to 0.5); (ii) model comparisons ba
Gene expression profiling by microarray technology has been successfully applied to classification and diagnostic prediction of cancers. Various machine learning and data mining methods are currently used for classifying gene expression data. However, these methods have not been developed to address the specific requirements of gene microarray analysis. First, microarray data is characterized by a high-dimensional feature space often exceeding the sample space dimensionality by a factor of 100 o
Commonly used accuracy-based performance values, with or without confidence intervals, are inadequate for comparing classifiers for small-sample data. We present a statistical methodology that avoids bias in cross-validated model selection in the context of small-sample scenarios. This methodology is valid for both k-fold cross-validation and repeated random sampling.
Deep learning in bioinformatics and biomedicine Daniel Berrar, Daniel Berrar Data Science Laboratory, Department of Information and Communications Engineering, Tokyo Institute of Technology, Ookayama, Tokyo 152-8550, Japan Corresponding author: Daniel Berrar, Data Science Laboratory, Department of Information and Communications Engineering, Tokyo Institute of Technology, Ookayama, Tokyo 152-8550, Japan. E-mail: daniel.berrar@ict.e.titech.ac.jp Search for other works by this author on: Oxford Aca
We present survival trees as an exploratory tool for revealing new insights into gene expression profiles in combination with clinical patient data. Survival trees partition the patient data studied into groups with similar survival outcomes and identify characteristic genetic profiles within these groups. We demonstrate the application of survival trees in a study involving the expression profiles of 3,588 genes in 211 lung adenocarcinoma patients. The survival tree identified a group of early-
Null hypothesis significance tests and their p-values currently dominate the statistical evaluation of classifiers in machine learning. Here, we discuss fundamental problems of this research practice. We focus on the problem of comparing multiple fully specified classifiers on a small-sample test set. On the basis of the method by Quesenberry and Hurst, we derive confidence intervals for the effect size, i.e. the difference in true classification performance. These confidence intervals disentang