정윤서 교수
Yoonseo Jung
고려대학교 통계학과 · 수학
연구실 소개
정윤서 교수의 연구실은 통계적 모델링과 머신러닝 기반의 고도화된 데이터 분석 기법을 중심으로, 특히 교차검증(Cross-Validation)의 신뢰성과 효율성을 향상시키는 데 초점을 맞추고 있습니다. 선형 모델 및 비선형 모델에서의 변수 선택, 추정 안정성 향상, 그리고 희소한 생물정보학 데이터(예: 유전자 다형성)에서의 정확한 추론을 위한 통계적 방법론을 개발하고 있습니다. 특히, 정규화된 분위수 회귀, 다중 SNP 동시 분석, 임상 데이터 기반의 모델 선택 기법 등 의료 및 유전체 연구에 응용 가능한 고도화된 통계 기법을 지속적으로 연구하고 있습니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15K-fold cross-validation (CV) is widely adopted as a model selection criterion. In K-fold CV, $ (K-1) $ (K−1) folds are used for model construction and the hold-out fold is allocated to model validation. This implies model construction is more emphasised than the model validation procedure. However, some studies have revealed that more emphasis on the validation procedure may result in improved model selection. Specifically, leave-m-out CV with n samples may achieve variable-selection consistency
Cross-validation (CV) type of methods have been widely used to facilitate model estimation and variable selection. In this work, we suggest a new K-fold CV procedure to select a candidate ‘optimal’ model from each hold-out fold and average the K candidate ‘optimal’ models to obtain the ultimate model. Due to the averaging effect, the variance of the proposed estimates can be significantly reduced. This new procedure results in more stable and efficient parameter estimation than the classical K-f
The Institute of Medicine has challenged oncology providers to address cancer survivorship care planning. Gaps in cancer survivorship knowledge are evident and will require focused education for this initiative to be successful.
Concomitant AF ablation in patients undergoing AVR resulted in increased sinus rhythm restoration, better echocardiographic results, and decreased anticoagulation requirement, without increasing surgical morbidity or mortality.
Quantile regression (QR) provides estimates of a range of conditional quantiles. This stands in contrast to traditional regression techniques, which focus on a single conditional mean function. Lee et al. [Regularization of case-specific parameters for robustness and efficiency. Statist Sci. 2012;27(3):350–372] proposed efficient QR by rounding the sharp corner of the loss. The main modification generally involves an asymmetric ℓ2 adjustment of the loss function around zero. We extend the idea o
In genome-wide association studies, the primary task is to detect biomarkers in the form of Single Nucleotide Polymorphisms (SNPs) that have nontrivial associations with a disease phenotype and some other important clinical/environmental factors. However, the extremely large number of SNPs comparing to the sample size inhibits application of classical methods such as the multiple logistic regression. Currently the most commonly used approach is still to analyze one SNP at a time. In this paper,
High-dimensional data are often encountered in biomedical, environmental, and other studies. For example, in biomedical studies that involve high-throughput omic data, an important problem is to search for genetic variables that are predictive of a particular phenotype. A conventional solution is to characterize such relationships through regression models in which a phenotype is treated as the response variable and the variables are treated as covariates; this approach becomes particularly chal
Outlying observations are often disregarded at the sacrifice of degrees of freedom or downsized via robust loss functions (e.g., Huber's loss) to reduce the undesirable impact on data analysis. In this article, we treat the outlying status of each observation as a parameter and propose a penalization method to automatically adjust the outliers. The proposed method shifts the outliers towards the fitted values, while preserve the non-outlying observations. We also develop a generally applicable a
The income or expenditure-related data sets are often nonlinear, heteroscedastic, skewed even after the transformation, and contain numerous outliers. We propose a class of robust nonlinear models that treat outlying observations effectively without removing them. For this purpose, case-specific parameters and a related penalty are employed to detect and modify the outliers systematically. We show how the existing nonlinear models such as smoothing splines and generalized additive models can be
The check loss function is used to define quantile regression. In cross-validation, it is also employed as a validation function when the true distribution is unknown. However, our empirical study indicates that validation with the check loss often leads to overfitting the data. In this work, we suggest a modified or L2-adjusted check loss which rounds the sharp corner in the middle of check loss. This has the effect of guarding against overfitting to some extent. The adjustment is devised to sh
As DNA microarray data contain relatively small<br> sample size compared to the number of genes, high dimensional<br> models are often employed. In high dimensional models, the selection<br> of tuning parameter (or, penalty parameter) is often one of the crucial<br> parts of the modeling. Cross-validation is one of the most common<br> methods for the tuning parameter selection, which selects a parameter<br> value with the smallest cross-validated score. However, selecting a<br> single value as a
Table Page 3.1 Difference in the number of selected variables for the fitted model to contaminated data from that to clean data . . . . . . . . . . . . . . .23 4.1 Point estimates and approximate 95% confidence intervals for MSE (multiplied by 1000), based on 200 replicates with n=300, and n=900, at selected quantiles. . . . . . . . . . . . . . . . . . . . . . . . . . . .56 5.1 Point estimates and approximate 95% confidence intervals for percentage reduction in mean MSE, based on 1000 replicates
대표 연구 분야
정윤서 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.