Seoul National University · Computer Science
Professor Yongdai Kim's research lab specializes in high-dimensional statistical modeling, with a focus on variable selection, regularization methods, and nonparametric Bayesian inference. The lab develops computationally efficient algorithms for sparse estimation in high-dimensional regression, such as SCAD and LASSO, and investigates their theoretical properties, including oracle and model selection consistency. It also explores Bayesian nonparametric methods for survival analysis and point processes, particularly using Lévy processes and neutral-to-the-right priors. The lab integrates statistical theory with applications in medical imaging and real-world data analysis.
Figures are computed from collected data and may differ slightly.
The smoothly clipped absolute deviation (SCAD) estimator, proposed by Fan and Li, has many desirable properties, including continuity, sparsity, and unbiasedness. The SCAD estimator also has the (asymptotically) oracle property when the dimension of covariates is fixed or diverges more slowly than the sample size. In this article we study the SCAD estimator in high-dimensional settings where the dimension of covariates can be much larger than the sample size. First, we develop an efficient optim
T2 weighted MR Image analysis of the paravertebral back muscles in patients with degenerative lumbar flat back showed significant fat infiltration compared with those in the normal control using digital image analysis. Digital image analysis of the paravertebral back muscles is a useful tool for measuring the degree of paravertebral back muscle degeneration.
Asymptotic properties of model selection criteria for high-dimensional regression models are studied where the dimension of covariates is much larger than the sample size. Several sufficient conditions for model selection consistency are provided. Non-Gaussian error distributions are considered and it is shown that the maximal number of covariates for model selection consistency depends on the tail behavior of the error distribution. Also, sufficient conditions for model selection consistency ar
LASSO (Least Absolute Shrinkage and Selection Operator) is a useful tool to achieve the shrinkage and variable selection simultaneously. Since LASSO uses the L1 penalty, the optimization should rely on the quadratic program (QP) or general non-linear program which is known to be computational intensive. In this paper, we propose a gradient descent algorithm for LASSO. Even though the final result is slightly less accurate, the proposed algorithm is computationally simpler than QP or non-linear p
This paper is concerned with nonparametric Bayesian inference of the Aalen’s multiplicative counting process model. For a desired nonparametric prior distribution of the cumulative intensity function, a class of Lévy processes is considered, and it is shown that the class of Lévy processes is conjugate for the multiplicative counting process model, and formulas for obtaining a posterior process are derived. Finally, our results are applied to several practically important models such as one poin
Ghosh and Ramamoorthi studied posterior consistency for survival models and showed that the posterior was consistent when the prior on the distribution of survival times was the Dirichlet process prior. In this paper,we study posterior consistency of survival models with neutral to the right process priors which include Dirichlet process priors. A set of sufficient conditions for posterior consistency with neutral to the right process priors are given. Interestingly, not all the neutral to the r
Considering that the transmission onset distribution peaked with the symptom onset and the pre-symptomatic transmission proportion is substantial, the usual preventive measures might be too late to prevent SARS-CoV-2 transmission.
In this study, the Delta variant of SARS-CoV-2 was estimated to propagate more easily among children and adolescents than pre-Delta strains, even after adjusting for contact pattern and vaccination status.
We propose two Bayesian bootstrap extensions, the binomial and Poisson forms, for proportional hazards models. The binomial form Bayesian bootstrap is the limit of the posterior distribution with a beta process prior as the amount of the prior information vanishes, and thus can be considered as a default nonparametric Bayesian analysis. It is also the same as Lo's Bayesian bootstrap for censored data when covariates are absent. The Poisson form Bayesian bootstrap is equivalent to the Bayesian an
빅데이터 시대를 맞이하여 통계학과 통계학자의 역할에 대하여 살펴본다. 빅데이터에 대한 정의 및 응용분야를 살펴보고, 빅데이터 자료의 통계학적 특징들 및 이와 관련한 통계학적 의의에 대해서 설명한다. 빅데이터 자료 분석에 유용하게 사용되는 통계적 방법론들에 대해서 살펴보고, 국외와 국내의 빅데이터 관련 프로젝트를 소개한다. We investigate the roles of statistics and statisticians in the big data era. Definition and application areas of big data are reviewed and statistical characteristics of big data and their meanings are discussed. Various statistical methodologies applicable to big data analysis are illustrated, and two real big data
Open papers in the app to read, cite, and organize with AI.