Jeongyoun Ahn
Korea Advanced Institute of Science and Technology · Computer Science
About the Lab
Professor Jeongyoun Ahn's research lab specializes in statistical methodology for complex, high-dimensional data, with a focus on functional data analysis, multivariate and interval-valued data modeling, and high-dimensional clustering and outlier detection. The lab develops innovative approaches to address challenges such as phase variation in functional data, batch effects in gene expression studies, and sample integrity in high-dimensional, low-sample-size (HDLSS) settings. A central theme is the integration of structural assumptions—such as sparsity, subspace geometry, and ordinal relationships—into robust statistical inference and machine learning frameworks. The lab’s work bridges theoretical statistics with practical applications in genomics, biomedicine, and data science.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15When functional data come as multiple curves per subject, characterizing the source of variations is not a trivial problem. The complexity of the problem goes deeper when there is phase variation in addition to amplitude variation. We consider clustering problem with multivariate functional data that have phase variations among the functional variables. We propose a conditional subject-specific warping framework in order to extract relevant features for clustering. Using multivariate growth curv
Abstract We consider interval‐valued data that frequently appear with advanced technologies in current data collection processes. Interval‐valued data refer to the data that are observed as ranges instead of single values. In the last decade, several approaches to the regression analysis of interval‐valued data have been introduced, but little work has been done on relevant statistical inferences concerning the regression model. In this paper, we propose a new approach to fit a linear regression
We propose a new hierarchical clustering method for high dimension, low sample size (HDLSS) data. The method utilizes the fact that each individ- ual data vector accounts for exactly one dimension in the subspace generated by HDLSS data. The linkage that is used for measuring the distance between clus- ters is the orthogonal distance between affine subspaces generated by each cluster. The ideal implementation would be to consider all possible binary splits of the data and choose the one that max
Batch bias has been found in many microarray gene expression studies that involve multiple batches of samples. A serious batch effect can alter not only the distribution of individual genes but also the inter-gene relationships. Even though some efforts have been made to remove such bias, there has been relatively less development on a multivariate approach, mainly because of the analytical difficulty due to the high-dimensional nature of gene expression data. We propose a multivariate batch adj
Despite the popularity of high dimension, low sample size data analysis, there has not been enough attention to the sample integrity issue, in particular, a possibility of outliers in the data. A new outlier detection procedure for data with much larger dimensionality than the sample size is presented. The proposed method is motivated by asymptotic properties of high-dimensional distance measures. Empirical studies suggest that high-dimensional outlier detection is more likely to suffer from a s
We consider the problem of fitting a generalized linear model with a three-dimensional image covariate, such as one obtained by functional magnetic resonance imaging (fMRI). A major challenge for fitting such a model is that the image is a multidimensional array, called a tensor, containing tens of thousands of elements, called voxels. Because there is a parameter associated with each voxel, fitting the model entails estimating tens of thousands of parameters with a typical sample size on the or
MOTIVATION: Ordinal classification problems arise in a variety of real-world applications, in which samples need to be classified into categories with a natural ordering. An example of classifying high-dimensional ordinal data is to use gene expressions to predict the ordinal drug response, which has been increasingly studied in pharmacogenetics. Classical ordinal classification methods are typically not able to tackle high-dimensional data and standard high-dimensional classification methods di
Abstract Kernel‐based classification methods, for example, support vector machines, map the data into a higher‐dimensional space via a kernel function. In practice, choosing the value of hyperparameter in the kernel function is crucial in order to ensure good performance. We propose a method of selecting the hyperparameter in the Gaussian radial basis function (RBF) kernel by considering the geometry of the embedded feature space. This method is independent of the choice of the discrimination al
Magnetic resonance imaging (MRI) is a clinically relevant, real-time imaging modality that is frequently utilized to assess stroke type and severity. However, specific MRI biomarkers that can be used to predict long-term functional recovery are still a critical need. Consequently, the present study sought to examine the prognostic value of commonly utilized MRI parameters to predict functional outcomes in a porcine model of ischemic stroke. Stroke was induced via permanent middle cerebral artery
This dissertation consists of three research topics regarding High Dimension, Low Sample Size (HDLSS) data analysis. The first topic is a study of the sample covariance matrix of a data set with extremely large dimensionality, but with relatively small sample size. Especially the asymptotic behavior of eigenvalues and eigenvectors of the sample covariance matrix is the focus of our study. Assuming that the true population covariance matrix of the data is not too far from identity matrix (i.e., s
In multi-class discrimination with high-dimensional data, identifying a lower-dimensional subspace with maximum class separation is crucial. We propose a new optimization criterion for finding such a discriminant subspace, which is the ratio of two traces: the trace of between-class scatter matrix and the trace of within-class scatter matrix. Since this problem is not well-defined for high-dimensional data, we propose to regularize the within trace and maximize the between trace. A careful inves
Gut microbiomes are increasingly found to be associated with many health-related characteristics of humans as well as animals. Regression with compositional microbiomes covariates is commonly used to identify important bacterial taxa that are related to various phenotype responses. Often the dimension of microbiome taxa easily exceeds the number of available samples, which creates a serious challenge in the estimation and inference of the model. The sparse log-contrast regression method is usefu
Research Areas
Dive deeper into Jeongyoun Ahn's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.