Yeunju Yoo
Seoul National University · Biochemistry, Genetics and Molecular Biology
About the Lab
Professor Yeunju Yoo's research lab specializes in statistical and computational methods for genetic data analysis, with a focus on linkage disequilibrium (LD) block construction, haplotype inference, and gene-based association testing in genome-wide association studies (GWAS). The lab develops innovative R packages such as gpart and methods like Big-LD to efficiently partition high-density SNP data into biologically meaningful blocks, improving statistical power and interpretability. Research also emphasizes addressing challenges in complex genomic regions—such as killer immunoglobulin-like receptor (KIR) genes—where genotype ambiguity complicates haplotype inference. The lab integrates population genetics theory with practical bioinformatics tools to enhance the analysis of next-generation sequencing data in complex disease studies.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15Motivation: Linkage disequilibrium (LD) block construction is required for research in population genetics and genetic epidemiology, including specification of sets of single nucleotide polymorphisms (SNPs) for analysis of multi-SNP based association and identification of haplotype blocks in high density sequencing data. Existing methods based on a narrow sense definition do not allow intermediate regions of low LD between strongly associated SNP pairs and tend to split high density SNP data int
SUMMARY: For the analysis of high-throughput genomic data produced by next-generation sequencing (NGS) technologies, researchers need to identify linkage disequilibrium (LD) structure in the genome. In this work, we developed an R package gpart which provides clustering algorithms to define LD blocks or analysis units consisting of SNPs. The visualization tool in gpart can display the LD structure and gene positions for up to 20 000 SNPs in one image. The gpart functions facilitate construction
MOTIVATION: Killer immunoglobulin-like receptor (KIR) genes vary considerably in their presence or absence on a specific regional haplotype. Because presence or absence of these genes is largely detected using locus-specific genotyping technology, the distinction between homozygosity and hemizygosity is often ambiguous. The performance of methods for haplotype inference (e.g. PL-EM, PHASE) for KIR genes may be compromised due to the large portion of ambiguous data. At the same time, many haploty
A central issue in genome-wide association (GWA) studies is assessing statistical significance while adjusting for multiple hypothesis testing. An equally important question is the statistical efficiency of the GWA design as compared to the traditional sequential approach in which genome-wide linkage analysis is followed by region-wise association mapping. Nevertheless, GWA is becoming more popular due in part to cost efficiency: commercially available 1M chips are nearly as inexpensive as a cus
By jointly analyzing multiple variants within a gene, instead of one at a time, gene-based multiple regression can improve power, robustness, and interpretation in genetic association analysis. We investigate multiple linear combination (MLC) test statistics for analysis of common variants under realistic trait models with linkage disequilibrium (LD) based on HapMap Asian haplotypes. MLC is a directional test that exploits LD structure in a gene to construct clusters of closely correlated varian
We performed a case-control association analysis of rheumatoid arthritis (RA) for several candidate genes using the North American Rheumatoid Arthritis Consortium (NARAC) data provided in Genetic Analysis Workshop 15. We conducted the case-control association analysis using all related cases and unrelated controls and compared the results with those from the analysis of samples using only one randomly selected case from each family and all unrelated controls. For both analyses we used a weighted
The power of genome-wide association studies can be improved by incorporating information from previous study findings, for example, results of genome-wide linkage analyses. Weighted false-discovery rate (FDR) control can incorporate genome-wide linkage scan results into the analysis of genome-wide association data by assigning single-nucleotide polymorphism (SNP) specific weights. Stratified FDR control can also be applied by stratifying the SNPs into high and low linkage strata. We applied the
Gene-based analysis of multiple single nucleotide polymorphisms (SNPs) in a gene region is an alternative to single SNP analysis. The multi-bin linear combination test (MLC) proposed in previous studies utilizes the correlation among SNPs within a gene to construct a gene-based global test. SNPs are partitioned into clusters of highly correlated SNPs, and the MLC test statistic quadratically combines linear combination statistics constructed for each cluster. The test has degrees of freedom equa
Multi-marker methods for genetic association analysis can be performed for common and low frequency SNPs to improve power. Regression models are an intuitive way to formulate multi-marker tests. In previous studies we evaluated regression-based multi-marker tests for common SNPs, and through identification of bins consisting of correlated SNPs, developed a multi-bin linear combination (MLC) test that is a compromise between a 1 df linear combination test and a multi-df global test. Bins of SNPs
2020년 1월 20일, 국내에서 코로나바이러스감염증-19의 첫 확진자가 발생하였다. 신종 감염병의 특성은 사회적으로 불안을 형성하였고, 필수 소비 품목으로 자리 잡은 마스크에 관한 각종 소비자 문제는 소비자불안을 촉발하는 계기가 되었다. 본 연구에서는 코로나19가 장기적인 영향을 미치는 가운데, 시기 및 상황의 변화에 따라 마스크에 관하여 소비자들이 어떠한 불안을 느끼는지 파악하고 이를 줄이기 위한 적절한 위험 커뮤니케이션 전략을 도출하고자 하였다. 이를 위해 트위터의 텍스트 데이터를 분석하였으며, 토픽모델링을 통해 시기별 주요 토픽을 추출하였다. 분석 결과, 마스크 관련 소비자불안에 대한 트윗의 버즈량 및 주요 키워드는 시기별로 차이가 있었다. 특히 집단감염이 발생하며 이른바 마스크 대란이 발생하였던 시기에 불안 관련 버즈량이 급증하여, 코로나19 상황에서 소비자불안에 가장 큰 영향을 준 요인은 마스크의 수급문제였음을 확인하였다. 또한, 토픽모델링을 실시하여 시기별 토픽과 그 내
상품큐레이션서비스의 등장 이래 시장의 성장세에도 불구하고 소비자들의 현실적인 목소리를 들어볼 수 있는 연구는 거의 이루어지지 않았다. 이에 본 연구에서는 기존 문헌을 토대로 상품큐레이션서비스를 ‘큐레이터가 선별하여구성한 상품을 소비자에게 제공하는 전자상거래 서비스’로 정의하고, 누가, 왜, 어떻게 서비스를 이용하는지 포괄적으로 탐색하고자 하였다. 이를 위해 심층면접을 실시하였으며 Glaser의 근거이론 방법론에 따라 자료를 분석한 결과 83개의 개념, 22개의 하위범주, 9개의 범주 및 ‘정보 과잉 환경에서 상품큐레이션서비스 이용을 자신에게 맞추어 나감’이라는 핵심범주가 도출되었다. 연구 결과, 상품큐레이션서비스 이용자들은 새로움과 효율성 추구 성향을 가지며 구매하는 상품에 대한 관여도가높았고, 정보 탐색에 대한 부담과 직접 선택에서의 한계와 더불어 서비스에 대한 다양한 기대로 서비스를 이용하는것으로 나타났다. 또한 이용자들은 서비스 이용 시 다양한 혜택과 문제를 지각했으며 다수는
The maximum LOD score statistic is extremely powerful for gene mapping when calculated using the correct genetic parameter value. When the mode of genetic transmission is unknown, the maximum of the LOD scores obtained using several genetic parameter values is reported. This latter statistic requires higher critical value than the maximum LOD score statistic calculated from a single genetic parameter value. In this paper, we compare the power of maximum LOD scores based on three fixed sets of ge
Over recent decades, machine learning, an integral subfield of artificial intelligence, has revolutionized diverse sectors, enabling data-driven decisions with minimal human intervention. In particular, the field of educational assessment emerges as a promising area for machine learning applications, where students can be classified and diagnosed using their performance data. The objectives of Diagnostic Classification Models (DCMs), which provide a suite of methods for diagnosing students' cogn
Research Areas
Dive deeper into Yeunju Yoo's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.