Pohang University of Science and Technology · Computer Science
Professor Hwanjo Yu's research lab specializes in data mining, machine learning, and knowledge discovery with a strong focus on privacy-preserving analytics, scalable learning algorithms, and biomedical data applications. The lab develops advanced methods for large-scale and distributed data mining, including privacy-preserving support vector machines and clustering-based learning techniques, while addressing real-world challenges in healthcare informatics through synthetic data generation and efficient literature search systems. Their work bridges theoretical rigor with practical deployment in domains such as electronic health records, high-dimensional omics data, and biomedical knowledge discovery.
Figures are computed from collected data and may differ slightly.
Support vector machines (SVMs) have been promising methods for classification and regression analysis because of their solid mathematical foundations which convery several salient properties that other methods hardly provide. However, despite the prominent properties of SVMs, they are not as favored for large-scale data mining as for pattern recognition or machine learning because the training complexity of SVMs is highly dependent on the size of a data set. Many real-world data mining applicati
Traditional Data Mining and Knowledge Discovery algorithms assume free access to data, either at a centralized location or in federated form. Increasingly, privacy and security concerns restrict this access, thus derailing data mining projects. What we need is distributed knowledge discovery that is sensitive to this problem. The key is to obtain valid results, while providing guarantees on the non-disclosure of data. Support vector machine classification is one of the most widely used classific
DAAE can effectively synthesize sequential EHRs by addressing its main challenges: the synthetic records should be realistic enough not to be distinguished from the real records, and they should cover all the training patients to reproduce the performance of specific downstream tasks.
Nowadays, high throughput experimental techniques make it feasible to examine and collect massive data at the molecular level. These data, typically mapped to a very high dimensional feature space, carry rich information about functionalities of certain chemical or biological entities and can be used to infer valuable knowledge for the purposes of classification and prediction. Typically, a small number of features or feature combinations may play determinant roles in functional discrimination.
Finding related articles from the PubMed (a large biomedical literature repository) is challenging because it is hard to express the user's specific relevance in the given query interface and a keyword query typically retrieves many results. Biomedical researchers spend a critical amount of time (e.g., often more than several days) in the literature search process. This paper proposes RefMed, a novel search system for PubMed, which supports relevance ranking by enabling relevance feedback on Pub
Open papers in the app to read, cite, and organize with AI.