Woncheol Jang
Seoul National University · 数学
研究室紹介
Professor Woncheol Jang's research lab specializes in statistical methodology for high-dimensional data analysis, with a focus on developing advanced regularization techniques for variable selection and clustering in the presence of correlated predictors. The lab investigates bias reduction in estimation procedures, particularly in generalized linear mixed models and case fatality rate estimation during disease outbreaks. A central theme is the development of novel penalized likelihood and constrained optimization methods—such as the generalized fused lasso and HORSES—that promote grouping of correlated variables while ensuring model interpretability and accuracy. The lab also applies these statistical innovations to real-world biomedical problems, including early detection of pancreatic cancer and epidemiological modeling of infectious diseases.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15The penalized quasi-likelihood (PQL) approach is the most common estimation procedure for the generalized linear mixed model (GLMM). However, it has been noticed that the PQL tends to underestimate variance components as well as regression coefficients in the previous literature. In this article, we numerically show that the biases of variance component estimates by PQL are systematically related to the biases of regression coefficient estimates by PQL, and also show that the biases of variance
BACKGROUND: Pancreatic cancer is the fourth leading cause of cancer-related deaths. Therefore, in order to improve survival rates, the development of biomarkers for early diagnosis is crucial. Recently, diabetes has been associated with an increased risk of pancreatic cancer. The aims of this study were to search for novel serum biomarkers that could be used for early diagnosis of pancreatic cancer and to identify whether diabetes was a risk factor for this disease. METHODS: Blood samples were c
Identifying homogeneous subgroups of variables can be challenging in high dimensional data analysis with highly correlated predictors. The generalized fused lasso has been proposed to simultaneously select correlated variables and identify them as predictive clusters (grouping property). In this article, we study properties of the generalized fused lasso. First, we present a geometric interpretation of the generalized fused lasso along with discussion of its persistency. Second, we analytically
This work is motivated by the recent worldwide pandemic of the novel coronavirus disease (COVID-19). When an epidemiological disease is prevalent, estimating the case fatality rate, the proportion of deaths out of the total cases, accurately and quickly is important as the case fatality rate is one of the crucial indicators of the risk of a disease. In this work, we propose an alternative estimator of the case fatality rate that provides more accurate estimate during an outbreak by reducing the
Identifying homogeneous subgroups of variables can be challenging in high dimensional data analysis with highly correlated predictors. We propose a new method called Hexagonal Operator for Regression with Shrinkage and Equality Selection, HORSES for short, that simultaneously selects positively correlated variables and identifies them as predictive clusters. This is achieved via a constrained least-squares problem with regularization that consists of a linear combination of an L_1 penalty for th
현대 과학기술의 발전으로 빅데이터의 시대가 도래하였다, 이러한 빅데이터는 여러가지 과학적 문제에 대한 해답을 제공하지만 반면에 이로 인해 새로운 도전에 직면하고 있다. 마이크로어레이 자료와 같은 고차원자료는 이러한 빅데이터에서 흔히 볼 수 있는 유형중의 하나이다. 본 논문에서는 고차원 자료분석에 많이 쓰이고 있는 대역검정과 동시검정, 그리고 이의 응용에 대한 소개를 한다. The power of modern technology is opening a new era of big data. The size of the datasets affords us the opportunity to answer many open scientific questions but also presents some interesting challenges. High-dimensional data such as microarray are common in big data. In this paper, we gi
We discuss nonparametric density estimation and regression for astrophysics problems. In particular, we show how to compute nonparametric confidence intervals for the location and size of peaks of a function. We illustrate these ideas with recent data on the Cosmic Microwave Background. We also briefly discuss nonparametric Bayesian inference.