임요한 교수
Johan Lim
서울대학교 첨단융합학부 · 수학
연구실 소개
임요한 교수의 연구실은 대량의 생물정보 데이터와 고차원 통계 데이터를 다루는 데 특화된 통계적 방법론 개발을 중심으로 활동하고 있습니다. 특히 유전자 배열 데이터 분석, 생존 분석에서의 비모수적 최대우도 추정, 고차원 제어도구 및 포트폴리오 최적화에 이르기까지, 복잡한 종속 구조나 차원의 저주 문제를 고려한 효과적인 통계 모델링을 연구합니다. 또한 전염병의 심각도 평가를 위한 실시간 사망률 추정 기법 등 실제 문제에 적용 가능한 실용적인 통계 기법 개발에도 기여하고 있습니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15In microarray data analysis, we are often required to combine several dependent partial test results. To overcome this, many suggestions have been made in previous literature; Tippett's test and Fisher's omnibus test are most popular. Both tests have known null distributions when the partial tests are independent. However, for dependent tests, their (even, asymptotic) null distributions are unknown and additional numerical procedures are required. In this paper, we revisited Stouffer's test base
The generalized T 2 chart (GT‐chart), which is composed of the T 2 statistic based on a small number of principal components and the remaining components, is a popular alternative to the traditional Hotelling's T 2 control chart. However, the application of the GT‐chart to high‐dimensional data, which are now ubiquitous, encounters difficulties from high dimensionality similar to other multivariate procedures. The sample principal components and their eigenvalues do not consistently estimate the
Many procedures have been proposed to compute the nonparametric maximum likelihood estimates (NPMLEs) of survival functions under various stochastic ordering constraints. Each of the existing procedures is applicable only to a specific type of stochastic order constraint and often hard to implement. In this paper, we describe a method for computing the NPMLEs of survival functions, based on geometric programming, that is applicable to more general constraints and easy to implement. To this end,
We study kernel density estimator from the ranked set samples (RSS). In the kernel density estimator, the selection of the bandwidth gives strong influence on the resulting estimate. In this article, we consider several different choices of the bandwidth and compare their asymptotic mean integrated square errors (MISE). We also propose a plug-in estimator of the bandwidth to minimize the asymptotic MISE. We numerically compare the MISE of the proposed kernel estimator (having the plug-in bandwid
We compare the performance of recently developed regularized covariance matrix estimators for Markowitz's portfolio optimization and of the minimum variance portfolio (MVP) problem in particular. We focus on seven estimators that are applied to the MVP problem in the literature; three regularize the eigenvalues of the sample covariance matrix, and the other four assume the sparsity of the true covariance matrix or its inverse. Comparisons are made with two sets of long-term S&P 500 stock return
This work is motivated by the recent Korean Middle East respiratory syndrome outbreak. We propose an easy online estimation procedure for the case fatality rate, ie, the proportion of deaths among the total cases during the course of an epidemic disease, which is an important indicator of the severity of a disease. The key step in our procedure is representing the data with the run-off triangle, which simultaneously takes into account two time axes, namely, the calendar and disease-duration time
Clinical researchers often analyze survival data as binary outcomes using the logistic regression method. This paper examines the information loss resulting from analyzing survival time as binary outcomes. We first demonstrate that, under the proportional hazard assumption, this binary discretization does result in a significant information loss. Second, when fitting a logistic model to survival time data, researchers inadvertently use the maximal statistic. We implement a numerical study to exa
Image segmentation is an important preprocessing step in a sophisticated and complex image processing algorithm. In segmenting real-world images, the cost of misclassification could depend on the true class. For example, in a two-class (negative or positive class) problem, the cost of misclassifying positive to negative class could not be equal to that of misclassifying negative to positive class. However, existing algorithms do not take into account the unequal misclassification cost. In this l
대표 연구 분야
임요한 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.