Skip to main content

Seunggeun Lee

Seoul National University · 生化学・遺伝学・分子生物学

研究室紹介

Professor Seunggeun Lee's research lab specializes in statistical genetics and computational biology, focusing on developing advanced statistical methods for large-scale genomic data analysis. The lab pioneers methods for rare variant association testing, polygenic risk score prediction across diverse ancestries, and efficient mixed model inference in biobank-scale data. Key research directions include improving type I error control in unbalanced case-control studies, correcting bias in principal component score prediction, and enhancing power through functional annotations and transfer learning. The lab's work bridges statistical methodology with practical applications in complex disease genetics and precision medicine.

statistical geneticspolygenic risk scoresrare variant analysisbiobank genomicsmixed models

Research Overview

Papers
164
Total Citations
11,650
Papers (5y)
54
Primary Field
生化学・遺伝学・分子生物学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
54total
2022
2023
2024
2025
2026
Citations per year (5y)
1,354total
20222023202420252026

Selected Papers

15
1
Article|1,575 citations·2018
Efficiently controlling for case-control imbalance and sample relatedness in large-scale genetic association studies
Wei Zhou, Jonas B. Nielsen, Lars G. Fritsche, Rounak Dey, Maiken E. Gabrielsen, Brooke N. Wolford, Jonathon LeFaive, Peter VandeHaar, Sarah A. Gagliano Taliun, Aliya Gifford, Lisa A. Bastarache, Wei‐Qi Wei
SJR Q1Nature Genetics
GeneticsBiochemistry, Genetics and Molecular Biology
2
Article|1,143 citations·2012
Optimal Unified Approach for Rare-Variant Association Testing with Application to Small-Sample Case-Control Whole-Exome Sequencing Studies
Seunggeun Lee, Mary J. Emond, Michael J. Bamshad, Kathleen C. Barnes, Mark J. Rieder, Deborah A. Nickerson, David C. Christiani, Mark M. Wurfel, Xihong Lin
SJR Q1The American Journal of Human GeneticsOA
GeneticsBiochemistry, Genetics and Molecular Biology
3
Article|259 citations·2013
General Framework for Meta-analysis of Rare Variants in Sequencing Association Studies
Seunggeun Lee, Tanya M. Teslovich, Michael Boehnke, Xihong Lin
SJR Q1The American Journal of Human GeneticsOA
GeneticsBiochemistry, Genetics and Molecular Biology
4
Article|162 citations·2017
A Fast and Accurate Algorithm to Test for Binary Phenotypes and Its Application to PheWAS
Rounak Dey, Ellen M. Schmidt, Gonçalo R. Abecasis, Seunggeun Lee
SJR Q1The American Journal of Human Genetics
GeneticsBiochemistry, Genetics and Molecular Biology
5
Article|141 citations·2022
SAIGE-GENE+ improves the efficiency and accuracy of set-based rare variant association tests
Wei Zhou, Wenjian Bi, Zhangchen Zhao, Kushal K. Dey, Karthik A. Jagadeesh, Konrad J. Karczewski, Mark J. Daly, Benjamin M. Neale, Seunggeun Lee
SJR Q1Nature GeneticsOA

Several biobanks, including UK Biobank (UKBB), are generating large-scale sequencing data. An existing method, SAIGE-GENE, performs well when testing variants with minor allele frequency (MAF) ≤ 1%, but inflation is observed in variance component set-based tests when restricting to variants with MAF ≤ 0.1% or 0.01%. Here, we propose SAIGE-GENE+ with greatly improved type I error control and computational efficiency to facilitate rare variant tests in large-scale data. We further show that incorp

GeneticsBiochemistry, Genetics and Molecular Biology
6
Preprint|128 citations·2017
Efficiently controlling for case-control imbalance and sample relatedness in large-scale genetic association studies
Wei Zhou, Jonas B. Nielsen, Lars G. Fritsche, Rounak Dey, Maiken E. Gabrielsen, Brooke N. Wolford, Jonathon LeFaive, Peter VandeHaar, Sarah A. Gagliano Taliun, Aliya Gifford, Lisa A. Bastarache, Wei‐Qi Wei
bioRxiv (Cold Spring Harbor Laboratory)OA

Abstract In genome-wide association studies (GWAS) for thousands of phenotypes in large biobanks, most binary traits have substantially fewer cases than controls. Both of the widely used approaches, linear mixed model and the recently proposed logistic mixed model, perform poorly – producing large type I error rates – in the analysis of phenotypes with unbalanced case-control ratios. Here we propose a scalable and accurate generalized mixed model association test that uses the saddlepoint approx

GeneticsBiochemistry, Genetics and Molecular Biology
7
Article|90 citations·2019
UK Biobank Whole-Exome Sequence Binary Phenome Analysis with Robust Region-Based Rare-Variant Test
Zhangchen Zhao, Wenjian Bi, Wei Zhou, Peter VandeHaar, Lars G. Fritsche, Seunggeun Lee
SJR Q1The American Journal of Human GeneticsOA
GeneticsBiochemistry, Genetics and Molecular Biology
8
Article|88 citations·2020
A Fast and Accurate Method for Genome-Wide Time-to-Event Data Analysis and Its Application to UK Biobank
Wenjian Bi, Lars G. Fritsche, Bhramar Mukherjee, Sehee Kim, Seunggeun Lee
SJR Q1The American Journal of Human GeneticsOA
GeneticsBiochemistry, Genetics and Molecular Biology
9
Article|87 citations·2010
Convergence and prediction of principal component scores in high-dimensional settings
Seunggeun Lee, Fei Zou, Fred A. Wright
SJR Q1The Annals of StatisticsOA

A number of settings arise in which it is of interest to predict Principal Component (PC) scores for new observations using data from an initial sample. In this paper, we demonstrate that naive approaches to PC score prediction can be substantially biased towards 0 in the analysis of large matrices. This phenomenon is largely related to known inconsistency results for sample eigenvalues and eigenvectors as both dimensions of the matrix increase. For the spiked eigenvalue model for random matrice

Statistics and ProbabilityMathematics
10
Article|83 citations·2022
The construction of cross-population polygenic risk scores using transfer learning
Zhangchen Zhao, Lars G. Fritsche, Jennifer A. Smith, Bhramar Mukherjee, Seunggeun Lee
SJR Q1The American Journal of Human GeneticsOA

As most existing genome-wide association studies (GWASs) were conducted in European-ancestry cohorts, and as the existing polygenic risk score (PRS) models have limited transferability across ancestry groups, PRS research on non-European-ancestry groups needs to make efficient use of available data until we attain large sample sizes across all ancestry groups. Here we propose a PRS method using transfer learning techniques. Our approach, TL-PRS, uses gradient descent to fine-tune the baseline PR

GeneticsBiochemistry, Genetics and Molecular Biology
11
Article|71 citations·2022
Genome-wide study on 72,298 individuals in Korean biobank data for 76 traits
Kisung Nam, Jangho Kim, Seunggeun Lee
SJR Q1Cell GenomicsOA

Genome-wide association studies (GWAS) on diverse ancestry groups are lacking, resulting in deficits of genetic discoveries and polygenic scores. We conducted GWAS for 76 phenotypes in Korean biobank data, namely the Korean Genome and Epidemiology Study (KoGES) (n = 72,298). Our analysis discovered 2,242 associated loci, including 122 novel associations, many of which were replicated in Biobank Japan (BBJ) GWAS. We also applied several up-to-date methods for genetic association tests to increase

GeneticsBiochemistry, Genetics and Molecular Biology
12
Article|70 citations·2018
Multi‐SKAT: General framework to test for rare‐variant association with multiple phenotypes
Diptavo Dutta, Laura J. Scott, Michael Boehnke, Seunggeun Lee
SJR Q2Genetic EpidemiologyOA

In genetic association analysis, a joint test of multiple distinct phenotypes can increase power to identify sets of trait-associated variants within genes or regions of interest. Existing multiphenotype tests for rare variants make specific assumptions about the patterns of association with underlying causal variants, and the violation of these assumptions can reduce power to detect association. Here, we develop a general framework for testing pleiotropic effects of rare variants on multiple co

GeneticsBiochemistry, Genetics and Molecular Biology
13
Article|61 citations·2015
An efficient resampling method for calibrating single and gene-based rare variant association analysis in case–control studies
Seunggeun Lee, Christian Fuchsberger, Sehee Kim, Laura J. Scott
SJR Q1BiostatisticsOA

For aggregation tests of genes or regions, the set of included variants often have small total minor allele counts (MACs), and this is particularly true when the most deleterious sets of variants are considered. When MAC is low, commonly used asymptotic tests are not well calibrated for binary phenotypes and can have conservative or anti-conservative results and potential power loss. Empirical p-values obtained via resampling methods are computationally costly for highly significant p-values and

GeneticsBiochemistry, Genetics and Molecular Biology
14
Article|58 citations·2020
Fast and robust ancestry prediction using principal component analysis
Daiwei Zhang, Rounak Dey, Seunggeun Lee
SJR Q1BioinformaticsOA

MOTIVATION: Population stratification (PS) is a major confounder in genome-wide association studies (GWAS) and can lead to false-positive associations. To adjust for PS, principal component analysis (PCA)-based ancestry prediction has been widely used. Simple projection (SP) based on principal component loadings and the recently developed data augmentation, decomposition and Procrustes (ADP) transformation, such as LASER and TRACE, are popular methods for predicting PC scores. However, the predi

GeneticsBiochemistry, Genetics and Molecular Biology
15
Preprint|39 citations·2019
Scalable generalized linear mixed model for region-based association tests in large biobanks and cohorts
Wei Zhou, Zhangchen Zhao, Jonas B. Nielsen, Lars G. Fritsche, Jonathon LeFaive, Sarah A. Gagliano Taliun, Wenjian Bi, Maiken E. Gabrielsen, Mark J. Daly, Benjamin M. Neale, Kristian Hveem, Gonçalo R. Abecasis
bioRxiv (Cold Spring Harbor Laboratory)OA

Abstract With very large sample sizes, population-based cohorts and biobanks provide an exciting opportunity to identify genetic components of complex traits. To analyze rare variants, gene or region-based multiple variant aggregate tests are commonly used to increase association test power. However, due to the substantial computation cost, existing region-based rare variant tests cannot analyze hundreds of thousands of samples while accounting for confounders, such as population stratification

GeneticsBiochemistry, Genetics and Molecular Biology

Research Areas

GeneticsMolecular BiologyStatistics and ProbabilityInfectious DiseasesPublic Health, Environmental and Occupational HealthExperimental and Cognitive Psychology

Seunggeun Leeの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。