Geunbaek Lee
Sungkyunkwan University · 情報科学
研究室紹介
Professor Geunbaek Lee's research lab specializes in statistical modeling for longitudinal and clustered categorical data, with a strong focus on developing advanced mixed-effects models that account for complex correlation structures and subject-specific variability. The lab emphasizes methodological innovations in random effects covariance matrix estimation, particularly through Bayesian approaches, Cholesky decomposition techniques, and noninformative priors that ensure proper model fitting under high-dimensional and positive-definite constraints. A key research direction involves improving the robustness and efficiency of marginalized random effects models and cumulative logit models by relaxing restrictive assumptions such as homogeneous correlation structures. The lab also contributes to theoretical statistics by deriving matching priors for generalized half-normal distributions, with applications in reliable inference for scale and shape parameters.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15Generalized linear mixed models(GLMMs) are frequently used for the analysis of longitudinal categorical data when the subject-specific effects is of interest. In GLMMs, the structure of the random effects covariance matrix is important for the estimation of fixed effects and to explain subject and time variations. The estimation of the matrix is not simple because of the high dimension and the positive definiteness; subsequently, we practically use the simple structure of the covariance matrix s
Marginalized random effects models (MREM) are commonly used toanalyze longitudinal categorical data when the population-averaged effects is ofinterest. In these models, random effects are used to explain both subject andtime variations. The estimation of the random effects covariance matrix is notsimple in MREM because of the high dimension and the positive definiteness. A relatively simple structure for the correlation is assumed such asa homogeneous AR(1) structure; however, it is too strong o
경시적 범주형자료 (longitudinal categorical data)는 의학, 보건학, 그리고 사회과학에서 많이 발생하는 자료이다. 이러한 자료는 반복측정으로 인한 결과치들의 상관관계를 설명하면서 공변량의 효과를 설명해야 한다. 이 논문에서 모집단에 대한 공변량의 효과를 추정하면서 우도함수에 기초한 모형인 주변화 변량효과모형 (marginalized random effects model)을 소개하고, 그 모형의 어떻게 발전했는지를 고찰한다. 그리고 실제 자료를 이용하여 제시된 모형을 설명한다.
Cumulative logit random effects models are typically used to analyze longitudinal ordinal data. The random effects covariance matrix is used in the models to demonstrate both subject-specific and time variations. The covariance matrix may also be homogeneous; however, the structure of the covariance matrix is assumed to be homoscedastic and restricted because the matrix is high-dimensional and should be positive definite. To satisfy these restrictions two Cholesky decomposition methods were prop
In this paper, we develop noninformative priors for the generalized half-normal distributionwhen scale and shape parameters are of interest, respectively. Especially, we developthe first and second order matching priors for both parameters. For the shape parameter,we reveal that the second order matching prior is a highest posterior density (HPD)matching prior and a cumulative distribution function (CDF) matching prior. In addition, itmatches the alternative coverage probabilities up to the seco
Longitudinal studies repeatedly measure outcomes over time. Therefore, repeated measurements are serially correlated from same subject (within-subject variation) and there is also variation between subjects (between-subject variation). The serial correlation and the between-subject variation must be taken into account to make proper inference on covariate effects (Diggle {\it et al.}, 2002). However, estimation of the covariance matrix is challenging because of many parameters and positive defin
경시적 자료분석에서 공변량 효과를 추정할 때 반복 측정된 결과들의 상관성은 고려되어야 한다. 따라서 공분산 행렬을 모형화하는 것은 매우 중요하다. 그러나 공분산 행렬의 추정은 모수들의 수가 많고 추정된 공분산행렬이 양정치성을 만족해야 하므로 쉽지 않은 문제이다. 이러한 제한을 극복하기 위해, 공분산행렬의 모형화를 위한 여러가지 방법을 제안하였다: 자기회귀/이동평균/자기회귀-이동평균 구조를 각각 적용한 수정콜레스키분해 (Pourahmadi, 1999), 이동평균 콜레스키분해 (Zhang과 Leng, 2012)와 자기회귀-이동평균 콜레스키 분해 (Lee 등, 2017) 이들 구조를 가지는 공분산 행렬의 특징을 비교연구하고자 한다. 이 세 가지 모형의 성능을 비교하기 위한 모의실험을 실시한다.
허들모형은 영이 과잉 가산자료를 분석하기 위해서 사용되어 왔다. 이 모형은 이산부분을 위한 로짓모형과 절삭된가산부분을 위한 절삭된 포아송모형의 혼합모형이다. 이 논문에서 우리는 경시적 영과잉 가산자료를 분석하기 위해서 수정된 콜레스키 분해을 이용하여 일반적인 이분산성을 가지는 변량효과 공분산행렬을 제안한다. 수정된 콜레스키 분해는 변량효과 공분산행렬을 일반화자기상관 모수와 혁신분산모수로 분리되면, 이러한 모수들은 베이지안 일반화 선형모형을 통해 추정된다. 그리고 실제 자료분석을 통하여 설명한다.
일반화 선형혼합모델은 일반적으로 경시적 범주형 자료를 분석하는데 사용된다. 이 모델에서 임의효과는 반복 측정치들의 시간에 따른 의존성을 설명한다. 임의효과 공분산행렬의 추정은 여러가지 제약조건들 때문에 쉽지 않은 문제이다. 제약조건으로는 행렬의 모수들의 수가 많으며, 또한 추정된 공분산행렬은 양정치성을 만족하여야 한다. 이러한 제한을 극복하기 위해, 임의효과 공분산행렬의 모형화를 위한 여러가지 방법이 제안되었다: 수정 쿌레스키분해, 이동평균 쿌레스키분해와 부분 자기상관행렬을 이용한 방법이 있다. 이 논문에서 위의 제안된 방법들을 소개한다.
Marginalized random effects models (MREMs) are often used to analyze longitudinal categorical data. The models permit direct estimation of marginal mean parameters and specify the serial correlation of longitudinal categorical data via the random effects. However, it is not easy to estimate the random effects covariance matrix in the MREMs because the matrix is high-dimensional and must be positive-definite. To solve these restrictions, we introduce two modeling approaches of the random effects
In longitudinal studies missing data are common and require a complicated analysis. There are two popular modeling frameworks, pattern mixture model (PMM) and selection models (SM) to analyze the missing data. We focus on the PMM and we also propose Bayesian pattern mixture models using generalized linear mixed models (GLMMs) for longitudinal binary data. Sensitivity analysis is used under the missing not at random assumption.
다변량 경시적 자료는 의학, 보건과학, 사회과학, 환경연구 등과 같은 많은 분야에서 측정된다. 이 자료는 시간에 따라 여러 개의 반응변수들이 반복적으로 측정되기 때문에 복잡한 상관관계를 가지고 있다. 즉 다른 시점에서의 동일한 반응변수들 간의 상관관계, 같은 시점에서의 서로 다른 반응변수들 간의 상관관계, 그리고 다른 시점에서의 서로 다른 반응변수들 간의 상관관계를 가지며, 이러한 복잡한 상관관계로 인해 다변량 경시적 자료에 대해 공분산행렬을 모형화하는 것은 단변량 경시적 자료분석에 비해 더 어렵다. 본 논문에서는 다변량 경시적 자료에 대한 공분산행렬을 모형화하는 것에 대한 여러 가지 접근법을 조사하고, 이 방법들 중에 해석이 용이한 Kim과 Zimmerman (2012)과 Lee 등 (2020)의 방법을 이용하여 실제 다변량 경시적 자료인 노동패널자료를 분석하고자 한다.
Modeling of the random effects covariance matrix in generalized linear mixed models (GLMMs) is an issue in analysis of longitudinal categorical data because the covariance matrix can be high-dimensional and its estimate must satisfy positive-definiteness. To satisfy these constraints, we consider the autoregressive and moving average Cholesky decomposition (ARMACD) to model the covariance matrix. The ARMACD creates a more flexible decomposition of the covariance matrix that provides generalized
다변량 경시적 자료에서 반복 측정된 자료들 사이에는 응답변수들 간의 세 가지 형태의 상관관계가 존재한다: 다른 시점에서 다른 반응변수들 간의 상관관계, 다른 시점에서의 동일한 반응변수들 간의 상관관계, 그리고 같은 시점에서의 반응변수들 간의 상관관계. 따라서 다변량 경시적 자료분석에서는 이러한 상관관계들을 모두 가지는 공분산행렬을 고려하여 모형화하는 것이 중요하다. 하지만 이러한 공분산행렬은 양정치성 (positive definiteness)을 만족해야 하고, 때로는 이분산성 (heterogeneous)을 가질 수 있다. 또한 반복 측정 횟수가 증가함에 따라 공분산행렬의 모수의 수는 기하급수적으로 증가하여 추정하기가 쉽지 않다. 이 어려움들을 해결하기 위해 자기회귀 (autoregressive) 구조, 자기회귀-이동평균 (autoregressive-moving average) 구조를 가지는 공분산 행렬의 모형화 방법이 제안되었다. Lee 등 (2020)과 Lee 등 (2019)은 다