Yoonseo Jung
Korea University · 数学
研究室紹介
Professor Yoonseo Jung's research lab specializes in statistical methodology with a focus on robust and efficient model selection, particularly in high-dimensional and complex data settings. The lab develops advanced cross-validation techniques and regularization methods for regression and variable selection, with applications in genomics, oncology, and clinical outcomes research. Key research directions include improving model stability through ensemble averaging in K-fold cross-validation, extending quantile regression to heterogeneous models, and integrating statistical learning with biomedical data to support precision medicine and survivorship care planning.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15K-fold cross-validation (CV) is widely adopted as a model selection criterion. In K-fold CV, $ (K-1) $ (K−1) folds are used for model construction and the hold-out fold is allocated to model validation. This implies model construction is more emphasised than the model validation procedure. However, some studies have revealed that more emphasis on the validation procedure may result in improved model selection. Specifically, leave-m-out CV with n samples may achieve variable-selection consistency
Cross-validation (CV) type of methods have been widely used to facilitate model estimation and variable selection. In this work, we suggest a new K-fold CV procedure to select a candidate ‘optimal’ model from each hold-out fold and average the K candidate ‘optimal’ models to obtain the ultimate model. Due to the averaging effect, the variance of the proposed estimates can be significantly reduced. This new procedure results in more stable and efficient parameter estimation than the classical K-f
The Institute of Medicine has challenged oncology providers to address cancer survivorship care planning. Gaps in cancer survivorship knowledge are evident and will require focused education for this initiative to be successful.
Concomitant AF ablation in patients undergoing AVR resulted in increased sinus rhythm restoration, better echocardiographic results, and decreased anticoagulation requirement, without increasing surgical morbidity or mortality.
Quantile regression (QR) provides estimates of a range of conditional quantiles. This stands in contrast to traditional regression techniques, which focus on a single conditional mean function. Lee et al. [Regularization of case-specific parameters for robustness and efficiency. Statist Sci. 2012;27(3):350–372] proposed efficient QR by rounding the sharp corner of the loss. The main modification generally involves an asymmetric ℓ2 adjustment of the loss function around zero. We extend the idea o
In genome-wide association studies, the primary task is to detect biomarkers in the form of Single Nucleotide Polymorphisms (SNPs) that have nontrivial associations with a disease phenotype and some other important clinical/environmental factors. However, the extremely large number of SNPs comparing to the sample size inhibits application of classical methods such as the multiple logistic regression. Currently the most commonly used approach is still to analyze one SNP at a time. In this paper,
High-dimensional data are often encountered in biomedical, environmental, and other studies. For example, in biomedical studies that involve high-throughput omic data, an important problem is to search for genetic variables that are predictive of a particular phenotype. A conventional solution is to characterize such relationships through regression models in which a phenotype is treated as the response variable and the variables are treated as covariates; this approach becomes particularly chal
Outlying observations are often disregarded at the sacrifice of degrees of freedom or downsized via robust loss functions (e.g., Huber's loss) to reduce the undesirable impact on data analysis. In this article, we treat the outlying status of each observation as a parameter and propose a penalization method to automatically adjust the outliers. The proposed method shifts the outliers towards the fitted values, while preserve the non-outlying observations. We also develop a generally applicable a
The income or expenditure-related data sets are often nonlinear, heteroscedastic, skewed even after the transformation, and contain numerous outliers. We propose a class of robust nonlinear models that treat outlying observations effectively without removing them. For this purpose, case-specific parameters and a related penalty are employed to detect and modify the outliers systematically. We show how the existing nonlinear models such as smoothing splines and generalized additive models can be
The check loss function is used to define quantile regression. In cross-validation, it is also employed as a validation function when the true distribution is unknown. However, our empirical study indicates that validation with the check loss often leads to overfitting the data. In this work, we suggest a modified or L2-adjusted check loss which rounds the sharp corner in the middle of check loss. This has the effect of guarding against overfitting to some extent. The adjustment is devised to sh
As DNA microarray data contain relatively small<br> sample size compared to the number of genes, high dimensional<br> models are often employed. In high dimensional models, the selection<br> of tuning parameter (or, penalty parameter) is often one of the crucial<br> parts of the modeling. Cross-validation is one of the most common<br> methods for the tuning parameter selection, which selects a parameter<br> value with the smallest cross-validated score. However, selecting a<br> single value as a
Table Page 3.1 Difference in the number of selected variables for the fitted model to contaminated data from that to clean data . . . . . . . . . . . . . . .23 4.1 Point estimates and approximate 95% confidence intervals for MSE (multiplied by 1000), based on 200 replicates with n=300, and n=900, at selected quantiles. . . . . . . . . . . . . . . . . . . . . . . . . . . .56 5.1 Point estimates and approximate 95% confidence intervals for percentage reduction in mean MSE, based on 1000 replicates