유연주 교수
Yeunju Yoo
서울대학교 · 생화학·유전·분자생물학
연구실 소개
유연주 교수의 연구실은 유전체 연관 분석(GWAS)의 통계적 유의성과 분석 설계의 효율성을 고도로 분석하며, 특히 다변량 유전자 기반 분석 기법과 연관된 통계적 방법론을 개발하고 있습니다. 특히 유전자 내 다수의 유전자 변이를 동시에 고려하는 MLC(다중 선형 조합) 검정법을 통해 연관성 탐색의 검정력과 해석 가능성 향상을 도모하고 있으며, 가족 간 유사성과 관련된 복잡한 설계에서도 정확한 유전자 연관 분석을 가능하게 하는 통계적 모델링에 주력하고 있습니다. 이는 류마티스성 관절염과 같은 복합질환의 유전적 기반 규명에 기여하고 있습니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15Supplementary data are available at Bioinformatics online.
Supplementary data are available at Bioinformatics online.
Supplementary data are available at Bioinformatics online.
A central issue in genome-wide association (GWA) studies is assessing statistical significance while adjusting for multiple hypothesis testing. An equally important question is the statistical efficiency of the GWA design as compared to the traditional sequential approach in which genome-wide linkage analysis is followed by region-wise association mapping. Nevertheless, GWA is becoming more popular due in part to cost efficiency: commercially available 1M chips are nearly as inexpensive as a cus
By jointly analyzing multiple variants within a gene, instead of one at a time, gene-based multiple regression can improve power, robustness, and interpretation in genetic association analysis. We investigate multiple linear combination (MLC) test statistics for analysis of common variants under realistic trait models with linkage disequilibrium (LD) based on HapMap Asian haplotypes. MLC is a directional test that exploits LD structure in a gene to construct clusters of closely correlated varian
We performed a case-control association analysis of rheumatoid arthritis (RA) for several candidate genes using the North American Rheumatoid Arthritis Consortium (NARAC) data provided in Genetic Analysis Workshop 15. We conducted the case-control association analysis using all related cases and unrelated controls and compared the results with those from the analysis of samples using only one randomly selected case from each family and all unrelated controls. For both analyses we used a weighted
The power of genome-wide association studies can be improved by incorporating information from previous study findings, for example, results of genome-wide linkage analyses. Weighted false-discovery rate (FDR) control can incorporate genome-wide linkage scan results into the analysis of genome-wide association data by assigning single-nucleotide polymorphism (SNP) specific weights. Stratified FDR control can also be applied by stratifying the SNPs into high and low linkage strata. We applied the
Gene-based analysis of multiple single nucleotide polymorphisms (SNPs) in a gene region is an alternative to single SNP analysis. The multi-bin linear combination test (MLC) proposed in previous studies utilizes the correlation among SNPs within a gene to construct a gene-based global test. SNPs are partitioned into clusters of highly correlated SNPs, and the MLC test statistic quadratically combines linear combination statistics constructed for each cluster. The test has degrees of freedom equa
Multi-marker methods for genetic association analysis can be performed for common and low frequency SNPs to improve power. Regression models are an intuitive way to formulate multi-marker tests. In previous studies we evaluated regression-based multi-marker tests for common SNPs, and through identification of bins consisting of correlated SNPs, developed a multi-bin linear combination (MLC) test that is a compromise between a 1 df linear combination test and a multi-df global test. Bins of SNPs
Over recent decades, machine learning, an integral subfield of artificial intelligence, has revolutionized diverse sectors, enabling data-driven decisions with minimal human intervention. In particular, the field of educational assessment emerges as a promising area for machine learning applications, where students can be classified and diagnosed using their performance data. The objectives of Diagnostic Classification Models (DCMs), which provide a suite of methods for diagnosing students' cogn
The maximum LOD score statistic is extremely powerful for gene mapping when calculated using the correct genetic parameter value. When the mode of genetic transmission is unknown, the maximum of the LOD scores obtained using several genetic parameter values is reported. This latter statistic requires higher critical value than the maximum LOD score statistic calculated from a single genetic parameter value. In this paper, we compare the power of maximum LOD scores based on three fixed sets of ge
상품큐레이션서비스의 등장 이래 시장의 성장세에도 불구하고 소비자들의 현실적인 목소리를 들어볼 수 있는 연구는 거의 이루어지지 않았다. 이에 본 연구에서는 기존 문헌을 토대로 상품큐레이션서비스를 ‘큐레이터가 선별하여 구성한 상품을 소비자에게 제공하는 전자상거래 서비스’로 정의하고, 누가, 왜, 어떻게 서비스를 이용하는지 포괄적으로 탐색하고자 하였다. 이를 위해 심층면접을 실시하였으며 Glaser의 근거이론 방법론에 따라 자료를 분석한 결과 83개의 개념, 22개의 하위범주, 9개의 범주 및 ‘정보 과잉 환경에서 상품큐레이션서비스 이용을 자신에게 맞추어 나감’이라는 핵심범주가 도출되었다. 연구 결과, 상품큐레이션서비스 이용자들은 새로움과 효율성 추구 성향을 가지며 구매하는 상품에 대한 관여도가 높았고, 정보 탐색에 대한 부담과 직접 선택에서의 한계와 더불어 서비스에 대한 다양한 기대로 서비스를 이용하는 것으로 나타났다. 또한 이용자들은 서비스 이용 시 다양한 혜택과 문제를 지각했으며 다
The power to detect linkage to the slope genes is quite low. But the power using disease-related traits as a phenotype is greater than the power using the disease (hypertension) phenotype.
본 연구는 가계부채가 지속적으로 증가하는 상황에서 2020년 가계금융복지조사 데이터를 활용하여 부채보유 가계의 특성을 파악하고 상환을 연체할 가능성이 높은 가계를 예측하였다. 소득계층별로 부채보유 가계를 유형화하였고, 선행연구에서 잘 다루어지지 않던 다양한 변수들을 활용함으로써 모형의 성능을 높였으며, 머신러닝의 대표적 방법인 의사결정나무 분석을 실시하였다는 점에서 기존의 연구들과 차별점을 가진다. 분석 결과, 소득이 증가할수록 연체 경험 비율과 주관적 부채부담 수준은 낮아지고 재무건전성은 양호해져, 저소득층의 취약성에 주목해야 할 필요성이 대두되었다. 또, 의사결정나무 분석 결과 소득계층을 막론하고 상환연체 가능성을 높이는 가장 중요한 변수는 주관적 부채부담이었으며 그 외의 변수는 소득계층에 따라 중요도가 상이한 것으로 나타났다. 저소득층과 중간소득층은 상환가능성과 DTA, 거주형태가 중요 변수였던 반면, 고소득층의 경우 신용카드 관련 대출 보유 여부, 상환가능성, 소비 목적 대출
대표 연구 분야
유연주 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.