Skip to main content

Jongsun Park

Sungkyunkwan University · Computer Science

About the Lab

Professor Jongsun Park's research lab specializes in statistical learning, high-dimensional data analysis, and advanced variable selection techniques. The lab focuses on developing innovative penalized regression and dimension reduction methods—such as nonconcave penalties, sparse principal component analysis, and sliced inverse regression with regularization—for improved variable selection and prediction accuracy in complex, high-dimensional datasets. The lab also integrates metaheuristic algorithms like genetic programming and particle swarm optimization with statistical models to enhance model interpretability and performance in classification and feature selection tasks. Their work bridges statistical theory with practical applications in finance, bioinformatics, and engineering.

variable selectionpenalized regressionhigh-dimensional datagenetic programmingMonte Carlo simulation

Research Overview

Papers
24
Total Citations
15
Papers (5y)
12
Primary Field
Computer Science

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
12total
2012
2013
2015
2017
2019
Citations per year (5y)
6total
20122013201520172019

Selected Papers

15
1
Article|4 citations·1998
Fisher consistency of GEE models under link misspecification
Chongsun Park, Sanford Weisberg
SJR Q1Computational Statistics & Data Analysis
Statistics and ProbabilityMathematics
2
Article|3 citations·2015
Variable Selection with Nonconcave Penalty Function on Reduced-Rank Regression
정상용, 박종선

In this article, we propose nonconcave penalties on a reduced-rankregression model to select variables and estimate coefficientssimultaneously. We apply HARD (hard thresholding) and SCAD(smoothly clipped absolute deviation) symmetric penalty functions with singularities at the origin, and bounded by aconstant to reduce bias. In our simulation study and real dataanalysis, the new method is compared with an existing variableselection method using L_1 penalty that exhibits competitiveperformance in

3
Article|2 citations·2017
Probabilistic penalized principal component analysis
Chongsun Park, Morgan C. Wang, Eun Bi Mo
SJR Q3Communications for Statistical Applications and MethodsOA

A variable selection method based on probabilistic principal component analysis (PCA) using penalized likelihood method is proposed. The proposed method is a two-step variable reduction method. The first step is based on the probabilistic principal component idea to identify principle components. The penalty function is used to identify important variables in each component. We then build a model on the original data space instead of building on the rotated data space through latent variables (p

Analytical ChemistryChemistry
4
Article|2 citations·2006
순차적으로 선택된 특성과 유전 프로그래밍을 이용한 결정나무
김효중, 박종선

Decision tree induction algorithm is one of the most widely used methods in classification problems. However, they could be trapped into a local minimum and have no reasonable means to escape from it if tree algorithm uses top-down search algorithm. Further, if irrelevant or redundant features are included in the data set, tree algorithms produces trees that are less accurate than those from the data set with only relevant features. We propose a hybrid algorithm to generate decision tree that u

5
Article|1 citations·2003
A Decision Tree Algorithm using Genetic Programming
Chongsun Park, Young Kyong Ko
SJR Q3Communications for Statistical Applications and MethodsOA

We explore the use of genetic programming to evolve decision trees directly for classification problems with both discrete and continuous predictors. We demonstrate that the derived hypotheses of standard algorithms can substantially deviated from the optimum. This deviation is partly due to their top-down style procedures. The performance of the system is measured on a set of real and simulated data sets and compared with the performance of well-known algorithms like CHAID, CART, C5.0, and QUES

Artificial IntelligenceComputer Science
6
Article|1 citations·2007
Variable Selection in Sliced Inverse Regression Using Generalized Eigenvalue Problem with Penalties
박종선

Variable selection algorithm for Sliced Inverse Regression using penaltyfunction is proposed. We noted SIR models can be expressed as generalizedeigenvalue decompositions and incorporated penalty functions on them. Wefound from small simulation that the HARD penalty function seems to bethe best in preserving original directions compared with other well-knownpenalty functions. Also it turned out to be eective in forcing coecientestimates zero for irrelevant predictors in regression analysis. Resu

7
Article|1 citations·2010
유전알고리즘을 이용한 최적 k-최근접이웃 분류기
박종선, 허균

분류분석에 사용되는 k-최근접이웃 분류기에 유전알고리즘을 적용하여 의미 있는 변수들과 이들에 대한 가중치 그리고 적절한 k를 동시에 선택하는 알고리즘을 제시하였다. 다양한 실제 자료에 대하여 기존의 여러 방법들과 교차타당성 방법을 통하여 비교한 결과 효과적인 것으로 나타났다.

8
Article|1 citations·2013
Application of quasi-Monte Carlo methods in multi-asset option pricing
Eun Bi Mo, Chongsun Park
Journal of the Korean Data and Information Science SocietyOA

본 연구에서는 다중자산 옵션 가격의 추정에 있어 자산의 수, 상관계수, 자산의 값들과 표준편차의 여러 조합에 대한 시뮬레이션을 통하여 저불일치 수열에 따르는 준난수 몬테칼로 방법들을 비교하였다. 결과적으로 준난수와 모로 역변환을 이용하는 것이 기본적인 몬테칼로 방법보다 정확하였으며 자산의 수와 관계없이 준난수 방법들 중 혼합법들이 더욱 효과적임을 알 수 있었다. Quasi-Monte Carlo method is known to have lower convergence rate than the standard Monte Carlo method. Quasi-Monte Carlo methods are using low discrepancy sequences as quasi-random numbers. They include Halton sequence, Faure sequence, and Sobol sequence. In this article, we compared standard Mo

Numerical AnalysisMathematics
9
Article|0 citations·1995
Subset Selection with Random Predictors
Chongsun Park, Sanford Weisberg
University of Minnesota Digital Conservancy (University of Minnesota)OA
Statistics and ProbabilityMathematics
10
Article|0 citations·2005
분할 역회귀모형에서 차원결정을 위한 점근검정법
박종선, 곽재근

As a promising technique for dimension reduction in regression analysis, SlicedInverse Regression (SIR) and an associated chi-square test for dimensionality were in-troduced by Li (1991). However, Li's test needs assumption of Normality for predictorsand found to be heavily dependent on the number of slices. We will provide a uniedasymptotic test for determining the dimensionality of the SIR model which is based on the probabilistic principal component analysis and free of normality assumption o

11
Article|0 citations·2004
Asymptotic Test for Dimensionality in Probabilistic Principal Component Analysis with Missing Values
박종선

In this talk we proposed an asymptotic test for dimensionality in the latent variable model for probabilistic principal component analysis with missing values at random. Proposed algorithm is a sequential likelihood ratio test for an appropriate Normal latent variable model for the principal component analysis. Modified EM-algorithm is used to find MLE for the model parameters. Results from simulations and real data sets give us promising evidences that the proposed method is useful in finding n

12
Article|0 citations·2008
벌점함수를 이용한 부분최소제곱 회귀모형에서의 변수선택
박종선, 문규종

본 논문에서는 반응변수가 하나 이상이고 설명변수들의 수가 관측치에 비하여 상대적으로 많은 경우에 널리 사용되는 부분최소제곱회귀모형에 벌점함수를 적용하여 모형에 필요한 설명변수들을 선택하는 문제를 고려하였다. 모형에 필요한 설명변수들은 각각의 잠재변수들에 대한 최적해 문제에 벌점함수를 추가한 후 모의담금질을 이용하여 선택하였다. 실제 자료에 대한 적용 결과 모형의 설명력 및 예측력을 크게 떨어뜨리지 않으면서 필요없는 변수들을 효과적으로 제거하는 것으로 나타나 부분최소제곱회귀모형에서 최적인 설명변수들의 부분집합을 선택하는데 적용될 수 있을 것이다.

13
Article|0 citations·2011
일반화추정방정식(GEE)에 대한 부스트랩의 적용
박종선, 전용문

본 논문에서는 일반화추정방정식(GEE)모형에 대한 부스트랩 방법의적용에 대하여 살펴본다. 다양한 부스트랩 방법들 중 GEE모형에 적용이가능한 잔차, 쌍 및 점수함수 부스트랩 방법을 가상 및 실제 자료들에적용한 결과 회귀계수들에 대한 추정치와 표준오차가 점근값들과차이를 보이는 것으로 나타났다. 따라서 표본수가 크지 않은 경우부스트랩 방법을 통하여 GEE모형에서의 회귀계수에 대한 추정치화표준편차를 구하는 것이 효과적임을 알 수 있다.

14
Article|0 citations·2019
하이브리드 드롭아웃
박종선, 이명규

수 많은 모수들을 가지고 있는 방대한 심층신경망은 매우 강력한 기계학습 방법이지만 모형의 과도한 융통성으로 인하여 과적합문제를 내포하고 있다. 드롭아웃 방법은 크기가 큰 신경망의 과적합 문제를 해결하는 다양한 방법들 중 하나이며 매우 효과적인 방법으로 알려져 있다. 드롭아웃 방법은 훈련과정에서 각각의 표본에 다른 모형을 적용하는데 이들 모형은 입력과 은닉층의 노드들을 무작위로 제거한 모형들 중에 임의로 선택된다. 본 연구에서는 임의로 선택된 모형에 둘 이상의 표본을 적용하여 모형의 가중치들에 대한 추정치의 안정성을 높이는 하이브리드 드롭아웃 방법을 제시하였다. 실제 자료를 이용한 시뮬레이션 결과 노드의 선택확률과 모형의 적합에 사용되는 표본의 수를 적절하게 선택하여 기존의 방법에 비하여 추정치의 변동성이 감소시킬 수 있었으며 동시에 검증자료에 대한 최저오차도 줄일 수 있음을 보였다.

15
Article|0 citations·2019
혼합회귀모형에서 콤포넌트 및 설명변수에 대한 벌점함수의 적용
박종선, 모은비

주어진 회귀자료에 유한혼합회귀모형을 적합하는 경우 적절한 성분의 수를 선택하고 선택된 각각의 회귀모형에서 의미있는 예측변수들의 집합을 선택하며 동시에 편의와 변동이 작은 회귀계수 추정치들을 얻는 것은 매우 중요하다. 본 연구에서는 혼합선형회귀모형에서 성분의 개수와 회귀계수에 벌점함수를 적용하여 적절한 성분의 수와 각 성분의 회귀모형에 필요한 설명변수들을 동시에 선택하는 방법을 제시하였다. 성분에 대한 벌점은 성분들의 로그값에 SCAD 벌점함수를 적용하였고 회귀계수들에는 SCAD와 더불어 MCP 및 Adplasso 벌점함수들을 사용하여 가상자료와 실제자료들에 대한 결과를 비교하였다. SCAD-SCAD 벌점함수 조합과 SCAD-MCP 조합의 경우 기존의 Luo 등 (2008)의 방법에서 문제가 되었던 과적합 문제를 해결함과 동시에 선택된 성분의 수와 회귀계수들을 효과적으로 선택하였으며 회귀계수들의 추정치에 대한 편의도 크지 않았다. 본 연구는 성분의 수가 알려져 있지 않은 회귀자료에서

Research Areas

Artificial IntelligenceStatistics and ProbabilityAnalytical ChemistryNumerical AnalysisControl and Systems Engineering

Dive deeper into Jongsun Park's research on Nubint

Open this lab's papers in the app to read with AI, summarize, and cite in your writing.