Hwanjo Yu
Pohang University of Science and Technology · Computer Science
About the Lab
Professor Hwanjo Yu's research lab specializes in scalable and privacy-preserving machine learning, with a strong focus on support vector machines (SVMs) and their applications in large-scale data mining and Web mining. The lab explores innovative methods such as clustering-based SVMs, positive example-based learning (PEBL), and privacy-preserving knowledge discovery to address challenges in data efficiency, scalability, and security. A key research direction involves developing active learning and selective sampling techniques for ranking functions, particularly in information retrieval contexts. The lab also emphasizes practical solutions for real-world data mining problems under constraints of data privacy and limited labeled examples.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15Web page classification is one of the essential techniques for Web mining. Specifically, classifying Web pages of a user-interesting class is the first step of mining interesting information from the Web. However, constructing a classifier for an interesting class requires laborious pre-processing such as collecting positive and negative training examples. For instance, in order to construct a homepage classifier, one needs to collect a sample of homepages (positive examples) and a sample of non
Web page classification is one of the essential techniques for Web mining because classifying Web pages of an interesting class is often the first step of mining the Web. However, constructing a classifier for an interesting class requires laborious preprocessing such as collecting positive and negative training examples. For instance, in order to construct a "homepage" classifier, one needs to collect a sample of homepages (positive examples) and a sample of nonhomepages (negative examples). In
Learning ranking (or preference) functions has been a major issue in the machine learning community and has produced many applications in information retrieval. SVMs (Support Vector Machines) - a classification and regression methodology - have also shown excellent performance in learning ranking functions. They effectively learn ranking functions of high generalization based on the large-margin principle and also systematically support nonlinear ranking by the kernel trick. In this paper, we pr
Support vector machines (SVMs) have been promising methods for classification and regression analysis because of their solid mathematical foundations which convery several salient properties that other methods hardly provide. However, despite the prominent properties of SVMs, they are not as favored for large-scale data mining as for pattern recognition or machine learning because the training complexity of SVMs is highly dependent on the size of a data set. Many real-world data mining applicati
Active sampling (also called active learning or selective sampling) has been extensively researched for classification and rank learning methods, which is to select the most informative samples from unlabeled data such that, once the samples are labeled, the accuracy of the function learned from the samples is maximized. While active sampling methods require learning a function at each iteration to find the most informative samples, this paper proposes passive sampling techniques for regression,
Most existing studies of text classification assume that the training data are completely labeled. In reality, however, many information retrieval problems can be more accurately described as learning a binary classifier from a set of incompletely labeled examples, where we typically have a small number of labeled positive examples and a very large number of unlabeled examples. In this paper, we study such a problem of performing Text Classification WithOut labeled Negative data TC-WON). In this
Adoption of Electronic Health Record (EHR) systems has led to collection of massive healthcare data, which creates oppor- tunities and challenges to study them. Computational phenotyping offers a promising way to convert the sparse and complex data into meaningful concepts that are interpretable to healthcare givers to make use of them. We propose a novel su- pervised nonnegative tensor factorization methodology that derives discriminative and distinct phenotypes. We represented co-occurrence of
The demonstration of diminished or scarred renal parenchyma in children is often the decisive factor in determining the future management of children with urinary tract malformations. Renal scintigraphy using technetium 99m-labelled dimercaptosuccinic acid (DMSA), computed tomography (CT) and intravenous urography (IU) were used to evaluate the renal parenchyma prior to ureter re-implantation in a series of 13 children. Their ages ranged from 5 months to 3 years 8 months. The indication for oper
Research Areas
Dive deeper into Hwanjo Yu's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.