Johan Lim
Seoul National University · 数学
研究室紹介
Professor Johan Lim's research lab specializes in statistical methodology with a focus on high-dimensional data analysis, survival analysis under stochastic constraints, and robust multivariate statistical process control. The lab develops innovative computational and inferential techniques for complex data structures, including microarray data, ranked set sampling, and epidemic time-series data, often integrating geometric programming and regularization methods. A recurring theme is the improvement of estimation accuracy and efficiency under structural assumptions such as monotonicity, sparsity, or symmetry.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15In microarray data analysis, we are often required to combine several dependent partial test results. To overcome this, many suggestions have been made in previous literature; Tippett's test and Fisher's omnibus test are most popular. Both tests have known null distributions when the partial tests are independent. However, for dependent tests, their (even, asymptotic) null distributions are unknown and additional numerical procedures are required. In this paper, we revisited Stouffer's test base
The generalized T 2 chart (GT‐chart), which is composed of the T 2 statistic based on a small number of principal components and the remaining components, is a popular alternative to the traditional Hotelling's T 2 control chart. However, the application of the GT‐chart to high‐dimensional data, which are now ubiquitous, encounters difficulties from high dimensionality similar to other multivariate procedures. The sample principal components and their eigenvalues do not consistently estimate the
Many procedures have been proposed to compute the nonparametric maximum likelihood estimates (NPMLEs) of survival functions under various stochastic ordering constraints. Each of the existing procedures is applicable only to a specific type of stochastic order constraint and often hard to implement. In this paper, we describe a method for computing the NPMLEs of survival functions, based on geometric programming, that is applicable to more general constraints and easy to implement. To this end,
We study kernel density estimator from the ranked set samples (RSS). In the kernel density estimator, the selection of the bandwidth gives strong influence on the resulting estimate. In this article, we consider several different choices of the bandwidth and compare their asymptotic mean integrated square errors (MISE). We also propose a plug-in estimator of the bandwidth to minimize the asymptotic MISE. We numerically compare the MISE of the proposed kernel estimator (having the plug-in bandwid
We compare the performance of recently developed regularized covariance matrix estimators for Markowitz's portfolio optimization and of the minimum variance portfolio (MVP) problem in particular. We focus on seven estimators that are applied to the MVP problem in the literature; three regularize the eigenvalues of the sample covariance matrix, and the other four assume the sparsity of the true covariance matrix or its inverse. Comparisons are made with two sets of long-term S&P 500 stock return
This work is motivated by the recent Korean Middle East respiratory syndrome outbreak. We propose an easy online estimation procedure for the case fatality rate, ie, the proportion of deaths among the total cases during the course of an epidemic disease, which is an important indicator of the severity of a disease. The key step in our procedure is representing the data with the run-off triangle, which simultaneously takes into account two time axes, namely, the calendar and disease-duration time
Clinical researchers often analyze survival data as binary outcomes using the logistic regression method. This paper examines the information loss resulting from analyzing survival time as binary outcomes. We first demonstrate that, under the proportional hazard assumption, this binary discretization does result in a significant information loss. Second, when fitting a logistic model to survival time data, researchers inadvertently use the maximal statistic. We implement a numerical study to exa
Image segmentation is an important preprocessing step in a sophisticated and complex image processing algorithm. In segmenting real-world images, the cost of misclassification could depend on the true class. For example, in a two-class (negative or positive class) problem, the cost of misclassifying positive to negative class could not be equal to that of misclassifying negative to positive class. However, existing algorithms do not take into account the unequal misclassification cost. In this l