Jongsun Park
Sungkyunkwan University · 情報科学
研究室紹介
Professor Jongsun Park's research lab specializes in statistical learning, high-dimensional data analysis, and advanced variable selection techniques. The lab focuses on developing innovative penalized regression and dimension reduction methods—such as nonconcave penalties, sparse principal component analysis, and sliced inverse regression with regularization—for improved variable selection and prediction accuracy in complex, high-dimensional datasets. The lab also integrates metaheuristic algorithms like genetic programming and particle swarm optimization with statistical models to enhance model interpretability and performance in classification and feature selection tasks. Their work bridges statistical theory with practical applications in finance, bioinformatics, and engineering.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15In this article, we propose nonconcave penalties on a reduced-rankregression model to select variables and estimate coefficientssimultaneously. We apply HARD (hard thresholding) and SCAD(smoothly clipped absolute deviation) symmetric penalty functions with singularities at the origin, and bounded by aconstant to reduce bias. In our simulation study and real dataanalysis, the new method is compared with an existing variableselection method using L_1 penalty that exhibits competitiveperformance in
A variable selection method based on probabilistic principal component analysis (PCA) using penalized likelihood method is proposed. The proposed method is a two-step variable reduction method. The first step is based on the probabilistic principal component idea to identify principle components. The penalty function is used to identify important variables in each component. We then build a model on the original data space instead of building on the rotated data space through latent variables (p
Decision tree induction algorithm is one of the most widely used methods in classification problems. However, they could be trapped into a local minimum and have no reasonable means to escape from it if tree algorithm uses top-down search algorithm. Further, if irrelevant or redundant features are included in the data set, tree algorithms produces trees that are less accurate than those from the data set with only relevant features. We propose a hybrid algorithm to generate decision tree that u
We explore the use of genetic programming to evolve decision trees directly for classification problems with both discrete and continuous predictors. We demonstrate that the derived hypotheses of standard algorithms can substantially deviated from the optimum. This deviation is partly due to their top-down style procedures. The performance of the system is measured on a set of real and simulated data sets and compared with the performance of well-known algorithms like CHAID, CART, C5.0, and QUES
Variable selection algorithm for Sliced Inverse Regression using penaltyfunction is proposed. We noted SIR models can be expressed as generalizedeigenvalue decompositions and incorporated penalty functions on them. Wefound from small simulation that the HARD penalty function seems to bethe best in preserving original directions compared with other well-knownpenalty functions. Also it turned out to be eective in forcing coecientestimates zero for irrelevant predictors in regression analysis. Resu
본 연구에서는 다중자산 옵션 가격의 추정에 있어 자산의 수, 상관계수, 자산의 값들과 표준편차의 여러 조합에 대한 시뮬레이션을 통하여 저불일치 수열에 따르는 준난수 몬테칼로 방법들을 비교하였다. 결과적으로 준난수와 모로 역변환을 이용하는 것이 기본적인 몬테칼로 방법보다 정확하였으며 자산의 수와 관계없이 준난수 방법들 중 혼합법들이 더욱 효과적임을 알 수 있었다. Quasi-Monte Carlo method is known to have lower convergence rate than the standard Monte Carlo method. Quasi-Monte Carlo methods are using low discrepancy sequences as quasi-random numbers. They include Halton sequence, Faure sequence, and Sobol sequence. In this article, we compared standard Mo
분류분석에 사용되는 k-최근접이웃 분류기에 유전알고리즘을 적용하여 의미 있는 변수들과 이들에 대한 가중치 그리고 적절한 k를 동시에 선택하는 알고리즘을 제시하였다. 다양한 실제 자료에 대하여 기존의 여러 방법들과 교차타당성 방법을 통하여 비교한 결과 효과적인 것으로 나타났다.
In this talk we proposed an asymptotic test for dimensionality in the latent variable model for probabilistic principal component analysis with missing values at random. Proposed algorithm is a sequential likelihood ratio test for an appropriate Normal latent variable model for the principal component analysis. Modified EM-algorithm is used to find MLE for the model parameters. Results from simulations and real data sets give us promising evidences that the proposed method is useful in finding n
멀티미디어의 발달은 고속 네트워크 발전과 이어졌으며, 모바일 장비의 성능 향상으로 높은 전송속도로 광대역 인터넷망 접속과 실내와 실외 사이의 끊김 없는 모바일 멀티캐스팅 서비스를 가능하게 했다. 멀티캐스팅 서비스는 모바일 멀티캐스팅을 기반으로 효율적인 그룹 통신을 지원한다. 그러나 모바일 멀티캐스팅 서비스는 터널 컨버젼스와 핸드오버 지연이라는 제약조건이 있다. 이와 같은 문제를 해결하기 위해 많은 프로토콜들이 연구되어 왔고, 핸드오버 방법 또한 연구대상이 되었다. PMIPv6 기반의 네트워크 에서 모바일 멀티캐스팅 서비스를 위한 도메인 간 최적화된 핸드오버 모델을 제안한다. 제안한 모델은 터널 컨버젼스를 제거하고 라우터 프로세싱을 줄여준다. 또한, 적합한 전송 메커니즘을 이용하여 빠른 핸드오버를 가능하게 한다. 또한 제안한 방식은 이전에 제안된 모델 보다 패킷 전송 비용과 지연시간이 감소하여 도메인 간에서 빠른 핸드오버를 보장한다.
As a promising technique for dimension reduction in regression analysis, SlicedInverse Regression (SIR) and an associated chi-square test for dimensionality were in-troduced by Li (1991). However, Li's test needs assumption of Normality for predictorsand found to be heavily dependent on the number of slices. We will provide a uniedasymptotic test for determining the dimensionality of the SIR model which is based on the probabilistic principal component analysis and free of normality assumption o
본 논문에서는 반응변수가 하나 이상이고 설명변수들의 수가 관측치에 비하여 상대적으로 많은 경우에 널리 사용되는 부분최소제곱회귀모형에 벌점함수를 적용하여 모형에 필요한 설명변수들을 선택하는 문제를 고려하였다. 모형에 필요한 설명변수들은 각각의 잠재변수들에 대한 최적해 문제에 벌점함수를 추가한 후 모의담금질을 이용하여 선택하였다. 실제 자료에 대한 적용 결과 모형의 설명력 및 예측력을 크게 떨어뜨리지 않으면서 필요없는 변수들을 효과적으로 제거하는 것으로 나타나 부분최소제곱회귀모형에서 최적인 설명변수들의 부분집합을 선택하는데 적용될 수 있을 것이다.
수 많은 모수들을 가지고 있는 방대한 심층신경망은 매우 강력한 기계학습 방법이지만 모형의 과도한 융통성으로 인하여 과적합문제를 내포하고 있다. 드롭아웃 방법은 크기가 큰 신경망의 과적합 문제를 해결하는 다양한 방법들 중 하나이며 매우 효과적인 방법으로 알려져 있다. 드롭아웃 방법은 훈련과정에서 각각의 표본에 다른 모형을 적용하는데 이들 모형은 입력과 은닉층의 노드들을 무작위로 제거한 모형들 중에 임의로 선택된다. 본 연구에서는 임의로 선택된 모형에 둘 이상의 표본을 적용하여 모형의 가중치들에 대한 추정치의 안정성을 높이는 하이브리드 드롭아웃 방법을 제시하였다. 실제 자료를 이용한 시뮬레이션 결과 노드의 선택확률과 모형의 적합에 사용되는 표본의 수를 적절하게 선택하여 기존의 방법에 비하여 추정치의 변동성이 감소시킬 수 있었으며 동시에 검증자료에 대한 최저오차도 줄일 수 있음을 보였다.