Seoul National University · 情報科学
Professor Sungroh Yoon's research lab specializes in computational and systems biology, focusing on transforming biomedical big data into actionable biological insights using advanced machine learning and bioinformatics approaches. The lab develops innovative computational tools for high-throughput analysis of omics data, biomedical signal processing (e.g., ECG and PPG for non-invasive blood pressure prediction), and microbiome analysis. Key research directions include deep learning applications in genomics and proteomics, automated analysis of capillary electrophoresis data, and understanding host-microbe interactions in diseases like atopic dermatitis. The lab emphasizes methodological innovation to enable scalable, accurate, and robust analysis of complex biological datasets.
Figures are computed from collected data and may differ slightly.
In the era of big data, transformation of biomedical big data into valuable knowledge has been one of the most important challenges in bioinformatics. Deep learning has advanced rapidly since the early 2000s and now demonstrates state-of-the-art performance in various fields. Accordingly, application of deep learning in bioinformatics to gain insight from data has been emphasized in both academia and industry. Here, we review deep learning in bioinformatics, presenting examples of current resear
A list of predicted modules is available from the authors upon request.
Data and codes are available in http://data.snu.ac.kr/pub/lncRNAnet.
Cardiovascular disease is the leading cause of death in the world. It is vital to prevent it by rapid diagnosis and appropriate management through periodic blood pressure (BP) measurement. Recently, many studies have been conducted on methods to measure BP without a cuff. One of the most common methods of predicting BP without a cuff is to use the correlation between pulse wave velocity (PWV) and BP. Studies that predict BP through PWV have two problems to overcome: 1) Additional efforts are req
We propose a computational method called high-throughput robust analysis for capillary electrophoresis (HiTRACE) to automate the key tasks in large-scale nucleic acid CE analysis, including the profile alignment that has heretofore been a rate-limiting step in the highest throughput experiments. We illustrate the application of HiTRACE on 13 datasets representing 4 different RNAs, 3 chemical modification strategies and up to 480 single mutant variants; the largest datasets each include 87 360 ba
The aim of this study was to evaluate changes in the skin surface microbiome in patients with atopic dermatitis during treatment. The effect of narrowband ultraviolet B phototherapy was also studied to determine the influence of exposure to ultraviolet. A total of 18 patients with atopic dermatitis were included in the study. Patients were divided into 2 groups based on treatment: 1 group treated with narrowband ultraviolet B phototherapy and topical corticosteroid, and the other group treated w
One of the most important advances in biology in recent years may be the discovery of RNAs that can regulate gene expression. As one kind of such functional noncoding RNAs, microRNAs (miRNAs) form a class of endogenous 19-23-nucleotide RNAs that can have important regulatory roles in animals and plants by targeting transcripts for cleavage or translational repression. Since the discovery of the very first miRNAs, computational methods have been an invaluable tool that can complement experimental
Drug metabolism is determined by the biochemical and physiological properties of the drug molecule. To improve the performance of a drug property prediction model, it is important to extract complex molecular dynamics from limited data. Recent machine learning or deep learning based models have employed the atom- and bond-type information, as well as the structural information to predict drug properties. However, many of these methods can be used only for the graph representations. Message passi
The biclustering method can be a very useful analysis tool when some genes have multiple functions and experimental conditions are diverse in gene expression measurement. This is because the biclustering approach, in contrast to the conventional clustering techniques, focuses on finding a subset of the genes and a subset of the experimental conditions that together exhibit coherent behavior. However, the biclustering problem is inherently intractable, and it is often computationally costly to fi
The codes and pre-trained models are available at https://github.com/mswzeus/TargetNet.
To demystify the "black box" property of deep neural networks for natural language processing (NLP), several methods have been proposed to interpret their predictions by measuring the change in prediction probability after erasing each token of an input. Since existing methods replace each token with a predefined value (i.e., zero), the resulting sentence lies out of the training data distribution, yielding misleading interpretations. In this study, we raise the out-of-distribution problem induc
Non-autoregressive neural machine translation (NART) models suffer from the multi-modality problem which causes translation inconsistency such as token repetition. Most recent approaches have attempted to solve this problem by implicitly modeling dependencies between outputs. In this paper, we introduce AligNART, which leverages full alignment information to explicitly reduce the modality of the target distribution. AligNART divides the machine translation task into (i) alignment estimation and
Our results demonstrate that we can cluster protein environments successfully using a simplified representation and K-means clustering algorithm. The rediscovery of known 3D motifs allows us to calibrate the size and intercluster distances that characterize useful clusters. This information will then allow us to find new clusters with similar characteristics that represent novel structural or functional sites.
Abstract We propose a multiscale approach to data integration that accounts for the varying resolving power of different data types from the very outset. Starting with a very coarse description, we match the production response at the wells by recursively refining the reservoir grid. A multiphase streamline simulator is utilized for modeling fluid flow in the reservoir. The well data is then integrated using conventional geostatistics, for example sequential simulation methods. There are several
Open papers in the app to read, cite, and organize with AI.