Skip to main content

송후성 교수

Hoseung Song

KAIST 산업및시스템공학과 · 생화학·유전·분자생물학

연구실 소개

송후성 교수의 연구실은 마이크로바이옴 데이터의 복잡한 특성—영점 과잉, 과분산, 비정규 분포—를 고려한 통계적 분석 기법을 개발하는 데 초점을 맞추고 있습니다. 특히 배치 효과 제거, 차이 분석, 군집 구조를 고려한 독립성 검정 등 마이크로바이옴 연구의 핵심 문제를 해결하기 위한 혁신적인 비모수적 및 커널 기반 방법론을 연구합니다. 또한 고차원, 비유럽거리 데이터에 적합한 변화점 탐지 및 분포 비교 기법도 함께 개발하여 생명의학 분야의 정밀의료 연구에 기여하고자 합니다.

마이크로바이옴배치 효과비모수적 통계커널 기반 검정군집 데이터

연구 현황

논문 수
21
총 인용 수
184
최근 5년 논문
21
주요 분야
생화학·유전·분자생물학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
21총합
2020
2021
2022
2023
2024
5개년 연도별 피인용 수
184총합
20202021202220232024

주요 논문

15
1
논문|인용수 129·2022
Batch effects removal for microbiome data via conditional quantile regression
Wodan Ling, Jiuyao Lu, Ni Zhao, Anju Lulla, Anna Plantinga, Weijia Fu, Angela Zhang, Hongjiao Liu, Hoseung Song, Zhigang Li, Jun Chen, Timothy W. Randolph
SJR Q1Nature CommunicationsOA

Batch effects in microbiome data arise from differential processing of specimens and can lead to spurious findings and obscure true signals. Strategies designed for genomic data to mitigate batch effects usually fail to address the zero-inflated and over-dispersed microbiome data. Most strategies tailored for microbiome data are restricted to association testing or specialized study designs, failing to allow other analytic goals or general designs. Here, we develop the Conditional Quantile Regre

Statistics and ProbabilityMathematics
2
논문|인용수 9·2021
Asymptotic distribution-free changepoint detection for data with repeated observations
Hoseung Song, Hao Chen
SJR Q1Biometrika

Summary A nonparametric framework for changepoint detection, based on scan statistics utilizing graphs that represent similarities among observations, is gaining attention owing to its flexibility and good performance for high-dimensional and non-Euclidean data sequences. However, this graph-based framework faces challenges when there are repeated observations in the sequence, which is often the case for discrete data such as network data. In this article we extend the graph-based framework to s

GeneticsBiochemistry, Genetics and Molecular Biology
3
논문|인용수 8·2023
Generalized kernel two-sample tests
Hoseung Song, Hao Chen
SJR Q1BiometrikaOA

Summary Kernel two-sample tests have been widely used for multivariate data to test equality of distributions. However, existing tests based on mapping distributions into a reproducing kernel Hilbert space mainly target specific alternatives and do not work well for some scenarios when the dimension of the data is moderate to high due to the curse of dimensionality. We propose a new test statistic that makes use of a common pattern under moderate and high dimensions and achieves substantial powe

Statistics and ProbabilityMathematics
4
논문|인용수 8·2023
Accommodating multiple potential normalizations in microbiome associations studies
Hoseung Song, Wodan Ling, Ni Zhao, Anna Plantinga, Courtney A. Broedlow, Nichole R. Klatt, Tiffany Hensley‐McBain, Michael C. Wu
SJR Q1BMC BioinformaticsOA

BACKGROUND: Microbial communities are known to be closely related to many diseases, such as obesity and HIV, and it is of interest to identify differentially abundant microbial species between two or more environments. Since the abundances or counts of microbial species usually have different scales and suffer from zero-inflation or over-dispersion, normalization is a critical step before conducting differential abundance analysis. Several normalization approaches have been proposed, but it is d

Molecular BiologyBiochemistry, Genetics and Molecular Biology
5
논문|인용수 6·2022
A fast kernel independence test for cluster-correlated data
Hoseung Song, Hongjiao Liu, Michael C. Wu
SJR Q1Scientific ReportsOA

Cluster-correlated data receives a lot of attention in biomedical and longitudinal studies and it is of interest to assess the generalized dependence between two multivariate variables under the cluster-correlated structure. The Hilbert-Schmidt independence criterion (HSIC) is a powerful kernel-based test statistic that captures various dependence between two random vectors and can be applied to an arbitrary non-Euclidean domain. However, the existing HSIC is not directly applicable to cluster-c

Food ScienceAgricultural and Biological Sciences
6
논문|인용수 6·2024
Association between vaginal microbiota and vaginal inflammatory immune markers in postmenopausal women
Elizabeth H. Byrne, Hoseung Song, Sujatha Srinivasan, David N. Fredricks, Susan D. Reed, Katherine A. Guthrie, Michael C. Wu, Caroline M. Mitchell
SJR Q1Menopause The Journal of The North American Menopause SocietyOA

OBJECTIVE: In premenopausal individuals, vaginal microbiota diversity and lack of Lactobacillus dominance are associated with greater mucosal inflammation, which is linked to a higher risk of cervical dysplasia and infections. It is not known if the association between the vaginal microbiota and inflammation is present after menopause, when the vaginal microbiota is generally higher-diversity and fewer people have Lactobacillus dominance. METHODS: This is a post hoc analysis of a subset of postm

MicrobiologyImmunology and Microbiology
7
dataset|인용수 3·2022
gTestsMulti: New Graph-Based Multi-Sample Tests
Hoseung Song, Hao Chen
OA

New multi-sample tests for testing whether multiple samples are from the same distribution. They work well particularly for high-dimensional data. Song, H. and Chen, H. (2022) &lt;<a href="https://doi.org/10.48550/arXiv.2205.13787" target="_top">doi:10.48550/arXiv.2205.13787</a>&gt;.

Artificial IntelligenceComputer Science
8
논문|인용수 3·2021
Preventing Failures by Dataset Shift Detection in Safety-Critical Graph Applications
Hoseung Song, Jayaraman J. Thiagarajan, Bhavya Kailkhura
SJR Q2Frontiers in Artificial IntelligenceOA

Dataset shift refers to the problem where the input data distribution may change over time (e.g., between training and test stages). Since this can be a critical bottleneck in several safety-critical applications such as healthcare, drug-discovery, etc., dataset shift detection has become an important research issue in machine learning. Though several existing efforts have focused on image/video data, applications with graph-structured data have not received sufficient attention. Therefore, in t

Artificial IntelligenceComputer Science
9
논문|인용수 3·2024
Practical and Powerful Kernel-Based Change-Point Detection
Hoseung Song, Hao Chen
SJR Q1IEEE Transactions on Signal Processing

Change-point analysis plays a significant role in various fields to reveal discrepancies in distribution in a sequence of observations. While a number of algorithms have been proposed for high-dimensional data, kernel-based methods have not been well explored due to difficulties in controlling false discoveries and mediocre performance. In this paper, we propose a new kernel-based framework that makes use of an important pattern of data in high dimensions to boost power. Analytic approximations

Control and Systems EngineeringEngineering
10
preprint|인용수 2·2022
New graph-based multi-sample tests for high-dimensional and non-Euclidean data
Hoseung Song, Hao Chen
arXiv (Cornell University)OA

Testing the equality in distributions of multiple samples is a common task in many fields. However, this problem for high-dimensional or non-Euclidean data has not been well explored. In this paper, we propose new nonparametric tests based on a similarity graph constructed on the pooled observations from multiple samples, and make use of both within-sample edges and between-sample edges, a straightforward but yet not explored idea. The new tests exhibit substantial power improvements over existi

Statistics and ProbabilityMathematics
11
dataset|인용수 2·2020
kerTests: Generalized Kernel Two-Sample Tests
Hoseung Song, Hao Chen
OA

New kernel-based test and fast tests for testing whether two samples are from the same distribution. They work well particularly for high-dimensional data. Song, H. and Chen, H. (2023) &lt;<a href="https://doi.org/10.48550/arXiv.2011.06127" target="_top">doi:10.48550/arXiv.2011.06127</a>&gt;.

Industrial and Manufacturing EngineeringEngineering
12
preprint|인용수 2·2020
Generalized Kernel Two-Sample Tests
Hoseung Song, Hao Chen
arXiv (Cornell University)OA

Kernel two-sample tests have been widely used for multivariate data to test equality of distributions. However, existing tests based on mapping distributions into a reproducing kernel Hilbert space mainly target specific alternatives and do not work well for some scenarios when the dimension of the data is moderate to high due to the curse of dimensionality. We propose a new test statistic that makes use of a common pattern under moderate and high dimensions and achieves substantial power improv

Statistics and ProbabilityMathematics
13
preprint|인용수 1·2020
Asymptotic distribution-free change-point detection for data with repeated observations
Hoseung Song, Hao Chen
arXiv (Cornell University)OA

In the regime of change-point detection, a nonparametric framework based on scan statistics utilizing graphs representing similarities among observations is gaining attention due to its flexibility and good performances for high-dimensional and non-Euclidean data sequences, which are ubiquitous in this big data era. However, this graph-based framework encounters problems when there are repeated observations in the sequence, which often happens for discrete data, such as network data. In this wor

Molecular BiologyBiochemistry, Genetics and Molecular Biology
14
preprint|인용수 1·2024
A robust, scalable K-statistic for quantifying immune cell clustering in spatial proteomics data
Julia Wrobel, Hoseung Song
arXiv (Cornell University)OA

Spatial summary statistics based on point process theory are widely used to quantify the spatial organization of cell populations in single-cell spatial proteomics data. Among these, Ripley's $K$ is a popular metric for assessing whether cells are spatially clustered or are randomly dispersed. However, the key assumption of spatial homogeneity is frequently violated in spatial proteomics data, leading to overestimates of cell clustering and colocalization. To address this, we propose a novel $K$

EpidemiologyMedicine
15
dataset|인용수 1·2022
kerSeg: New Kernel-Based Change-Point Detection
Hoseung Song, Hao Chen
OA

New kernel-based test and fast tests for detecting change-points or changed-intervals where the distributions abruptly change. They work well particularly for high-dimensional data. Song, H. and Chen, H. (2022) &lt;<a href="https://doi.org/10.48550/arXiv.2206.01853" target="_top">doi:10.48550/arXiv.2206.01853</a>&gt;.

Control and Systems EngineeringEngineering

대표 연구 분야

Statistics and ProbabilityArtificial IntelligenceGeneticsMolecular BiologyControl and Systems EngineeringFood Science

송후성 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.