장원철 교수
Woncheol Jang
서울대학교 통계학과 · 수학
연구실 소개
장원철 교수의 연구실은 고차원 데이터에서 상관관계가 높은 변수들을 효과적으로 그룹화하고 선별하는 데 중점을 두고 있으며, 특히 일반화된 융합 로지스틱 회귀(Generalized Fused Lasso)와 유사한 정규화 기법을 활용한 변수 선택 및 그룹화 기법을 개발하고 있습니다. 또한, 생물의학 분야에서의 응용을 고려해, 당뇨병과 췌장암 간의 연관성 탐색, 코로나19의 사망률 추정 등 실제 임상 및 공중보건 문제에 적용 가능한 통계적 모델링 기법을 개발하고 있습니다. 특히, 페널티 최대우도법(PQL)의 편향 문제나 사망률 추정의 부정확성 등 실용적 통계 문제에 대한 이론적 분석과 개선 방안을 제시하는 데도 기여하고 있습니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15The penalized quasi-likelihood (PQL) approach is the most common estimation procedure for the generalized linear mixed model (GLMM). However, it has been noticed that the PQL tends to underestimate variance components as well as regression coefficients in the previous literature. In this article, we numerically show that the biases of variance component estimates by PQL are systematically related to the biases of regression coefficient estimates by PQL, and also show that the biases of variance
BACKGROUND: Pancreatic cancer is the fourth leading cause of cancer-related deaths. Therefore, in order to improve survival rates, the development of biomarkers for early diagnosis is crucial. Recently, diabetes has been associated with an increased risk of pancreatic cancer. The aims of this study were to search for novel serum biomarkers that could be used for early diagnosis of pancreatic cancer and to identify whether diabetes was a risk factor for this disease. METHODS: Blood samples were c
Identifying homogeneous subgroups of variables can be challenging in high dimensional data analysis with highly correlated predictors. The generalized fused lasso has been proposed to simultaneously select correlated variables and identify them as predictive clusters (grouping property). In this article, we study properties of the generalized fused lasso. First, we present a geometric interpretation of the generalized fused lasso along with discussion of its persistency. Second, we analytically
This work is motivated by the recent worldwide pandemic of the novel coronavirus disease (COVID-19). When an epidemiological disease is prevalent, estimating the case fatality rate, the proportion of deaths out of the total cases, accurately and quickly is important as the case fatality rate is one of the crucial indicators of the risk of a disease. In this work, we propose an alternative estimator of the case fatality rate that provides more accurate estimate during an outbreak by reducing the
Identifying homogeneous subgroups of variables can be challenging in high dimensional data analysis with highly correlated predictors. We propose a new method called Hexagonal Operator for Regression with Shrinkage and Equality Selection, HORSES for short, that simultaneously selects positively correlated variables and identifies them as predictive clusters. This is achieved via a constrained least-squares problem with regularization that consists of a linear combination of an L_1 penalty for th
현대 과학기술의 발전으로 빅데이터의 시대가 도래하였다, 이러한 빅데이터는 여러가지 과학적 문제에 대한 해답을 제공하지만 반면에 이로 인해 새로운 도전에 직면하고 있다. 마이크로어레이 자료와 같은 고차원자료는 이러한 빅데이터에서 흔히 볼 수 있는 유형중의 하나이다. 본 논문에서는 고차원 자료분석에 많이 쓰이고 있는 대역검정과 동시검정, 그리고 이의 응용에 대한 소개를 한다. The power of modern technology is opening a new era of big data. The size of the datasets affords us the opportunity to answer many open scientific questions but also presents some interesting challenges. High-dimensional data such as microarray are common in big data. In this paper, we gi
We discuss nonparametric density estimation and regression for astrophysics problems. In particular, we show how to compute nonparametric confidence intervals for the location and size of peaks of a function. We illustrate these ideas with recent data on the Cosmic Microwave Background. We also briefly discuss nonparametric Bayesian inference.
대표 연구 분야
장원철 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.