Wooju Lee
Seoul National University · Medicine
About the Lab
Professor Wooju Lee's research lab specializes in statistical methodology and data science applications in health and biomedical research, with a strong focus on developing and refining analytical techniques for complex, high-dimensional, and dependent data structures. The lab's work spans statistical inference in diagnostic test evaluation, machine learning for clinical prediction modeling in cardiovascular disease, sparse and regularized methods for genomics, and collective risk modeling in actuarial science. A recurring theme is the development of robust, interpretable, and computationally efficient statistical models that address real-world challenges in medicine, public health, and biostatistics.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15McNemar's test is often used in practice to compare the sensitivities and specificities for the evaluation of two diagnostic tests. For correct evaluation of accuracy, an intuitive recommendation is to test the diseased and the non-diseased groups separately so that the sensitivities can be compared among the diseased, and specificities can be compared among the healthy group of people. This paper provides a rigorous theoretical framework for this argument and study the validity of McNemar's tes
Machine learning (ML) has been suggested to improve the performance of prediction models. Nevertheless, research on predicting the risk in patients with acute myocardial infarction (AMI) has been limited and showed inconsistency in the performance of ML models versus traditional models (TMs). This study developed ML-based models (logistic regression with regularization, random forest, support vector machine, and extreme gradient boosting) and compared their performance in predicting the short- a
Canonical covariance analysis (CCA) has gained popularity as a method for the analysis of two sets of high-dimensional genomic data. However, it is often difficult to interpret the results because canonical vectors are linear combinations of all variables, and the coefficients are typically nonzero. Several sparse CCA methods have recently been proposed for reducing the number of nonzero coefficients, but these existing methods are not satisfactory because they still give too many nonzero coeffi
SOX11 is a transcription factor that is normally expressed in the fetal brain and has also been detected in some malignant tumors, including mantle cell lymphoma (MCL). MCL is a mature B-cell lymphoma that characteristically expresses cyclin D1, which has been used as a diagnostic tumor marker. SOX11 has also recently emerged as a tumor marker for MCL, particularly in cyclin D1-negative MCLs and to distinguish between MCLs and other cyclin D1-positive lymphomas. In this study, we evaluated the d
Several collective risk models have recently been proposed by relaxing the widely used but controversial assumption of independence between claim frequency and severity. Approaches include the bivariate copula model, random effect model, and two-part frequency-severity model. This study focuses on the copula approach to develop collective risk models that allow a flexible dependence structure for frequency and severity. We first revisit the bivariate copula method for frequency and average sever
The COVID-19 pandemic has brought about valuable insights regarding models, data, and experiments. In this narrative review, we summarised the existing literature on these three themes, exploring the challenges of providing forecasts, the requirement for real-time linkage of health-related datasets, and the role of 'experimentation' in evaluating interventions. This literature review encourages us to broaden our perspective for the future, acknowledging the significance of investing in models, d
The partial least-square (PLS) method has been adapted to the Cox's proportional hazards model for analyzing high-dimensional survival data. But because the latent components constructed in PLS employ all predictors regardless of their relevance, it is often difficult to interpret the results. In this paper, we propose a new formulation of sparse PLS (SPLS) procedure for survival data to allow simultaneous sparse variable selection and dimension reduction. We develop a computing algorithm for SP
Research Areas
Dive deeper into Wooju Lee's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.