Sung-jun Shin
Korea University · 生化学・遺伝学・分子生物学
研究室紹介
Professor Sung-jun Shin's research lab specializes in statistical genetics and high-dimensional data analysis, with a focus on developing advanced statistical models for genetic risk prediction and penetrance estimation in hereditary cancer syndromes such as Li-Fraumeni syndrome. The lab integrates family pedigree structures and competing risks models to correct for ascertainment bias and improve clinical characterization of high-risk individuals. Key methodological contributions include Bayesian semiparametric models, family-wise likelihoods, and novel sufficient dimension reduction techniques tailored for binary classification and survival analysis in complex genetic data.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15Abstract Li-Fraumeni syndrome (LFS) is a rare hereditary cancer syndrome associated with an autosomal-dominant mutation inheritance in the TP53 tumor suppressor gene and a wide spectrum of cancer diagnoses. The previously developed R package, LFSPRO, is capable of estimating the risk of an individual being a TP53 mutation carrier. However, an accurate estimation of the penetrance of different cancer types in LFS is crucial to improve the clinical characterization and management of high-risk indi
In statistics, data are not regarded as just numbers, but realizations of random elements. Suppose we have data y∼P . Then data analysis is essentially the process of uncovering the data generating...
Sufficient dimension reduction is popular for reducing data dimensionality without stringent model assumptions. However, most existing methods may work poorly for binary classification. For example, sliced inverse regression (Li, 1991) can estimate at most one direction if the response is binary. In this paper we propose principal weighted support vector machines, a unified framework for linear and nonlinear sufficient dimension reduction in binary classification. Its asymptotic properties are s
In high-dimensional data analysis, it is of primary interest to reduce the data dimensionality without loss of information. Sufficient dimension reduction (SDR) arises in this context, and many successful SDR methods have been developed since the introduction of sliced inverse regression (SIR) [Li (1991) Journal of the American Statistical Association 86, 316-327]. Despite their fast progress, though, most existing methods target on regression problems with a continuous response. For binary clas
Abstract Li-Fraumeni syndrome (LFS) is a rare autosomal dominant disorder associated with TP53 germline mutations and an increased lifetime risk of multiple primary cancers (MPC). Penetrance estimation of time to first and second primary cancer within LFS remains challenging because of limited data and the difficulty of characterizing the effects of a primary cancer on the penetrance of a second primary cancer. Using a recurrent events survival modeling approach that incorporates a family-wise l
Penetrance, which plays a key role in genetic research, is defined as the proportion of individuals with the genetic variants (i.e., genotype) that cause a particular trait and who have clinical symptoms of the trait (i.e., phenotype). We propose a Bayesian semiparametric approach to estimate the cancer-specific age-at-onset penetrance in the presence of the competing risk of multiple cancers. We employ a Bayesian semiparametric competing risk model to model the duration until individuals in a h
A common phenomenon in cancer syndromes is for an individual to have multiple primary cancers (MPC) at different sites during his/her lifetime. Patients with Li-Fraumeni syndrome (LFS), a rare pediatric cancer syndrome mainly caused by germline TP53 mutations, are known to have a higher probability of developing a second primary cancer than those with other cancer syndromes. In this context, it is desirable to model the development of MPC to enable better clinical management of LFS. Here, we pro
The support vector machine (SVM) is a popular learning method for binary classification. Standard SVMs treat all the data points equally, but in some practical problems it is more natural to assign different weights to observations from different classes. This leads to a broader class of learning, the so-called weighted SVMs (WSVMs), and one of their important applications is to estimate class probabilities besides learning the classification boundary. There are two parameters associated with th
Change‐point detection regains much attention recently for analyzing array or sequencing data for copy number variation (CNV) detection. In such applications, the true signals are typically very short and buried in the long data sequence, which makes it challenging to identify the variations efficiently and accurately. In this article, we propose a new change‐point detection method, a backward procedure, which is not only fast and simple enough to exploit high‐dimensional data but also performs