Korea Advanced Institute of Science and Technology · 情報科学
Professor Sungwon Han's research lab specializes in advancing data-driven solutions for societal challenges through innovative machine learning and AI techniques. The lab focuses on developing reliable methods for economic development measurement using satellite imagery, advancing federated learning with robust defense mechanisms against data poisoning, and enhancing fairness in machine learning through self-supervised representation learning. Key research directions include AI for social good, trustworthy and fair AI, and scalable intelligent systems for real-world applications.
Figures are computed from collected data and may differ slightly.
Reliable and timely measurements of economic activities are fundamental for understanding economic development and designing government policies. However, many developing countries still lack reliable data. In this paper, we introduce a novel approach for measuring economic development from high-resolution satellite images in the absence of ground truth statistics. Our method consists of three steps. First, we run a clustering algorithm on satellite images that distinguishes artifacts from natur
Satellite imagery has long been an attractive data source providing a wealth of information regarding human-inhabited areas. While high-resolution satellite images are rapidly becoming available, limited studies have focused on how to extract meaningful information regarding human habitation patterns and economic scales from such data. We present READ, a new approach for obtaining essential spatial representation for any given district from high-resolution satellite imagery based on deep neural
Federated learning is used to train a shared model in a decentralized way without clients sharing private data with each other. Federated learning systems are susceptible to poisoning attacks when malicious clients send false updates to the central server. Existing defense strategies are ineffective under non-IID data settings. This paper proposes a new defense strategy, FedCPA (Federated learning with Critical Parameter Analysis). Our attack-tolerant aggregation method is based on the observati
Two types of topic modeling predominate: generative methods that employ probabilistic latent models and clustering methods that identify semantically coherent groups. This paper newly presents UTopic (Unified neural Topic model via contrastive learning and term weighting) that combines the advantages of these two types. UTopic uses contrastive learning and term weighting to learn knowledge from a pretrained language model and discover influential terms from semantically coherent clusters. Experi
The Internet of Things (IoT) has recently become mainstream, different from the incomplete realization of ubiquitous computing, sensor networks, and others in past that commonly shared the notion of visionary hyper-connectivity, i.e., everything is connected. Among many, one significant reason for this achievement of IoT is the nowadays availability of cloud computing which can cost effectively deal with the sheer number of globally distributed devices and data generated from those devices. In t
Algorithmic fairness has become an important machine learning problem, especially for mission-critical Web applications. This work presents a self-supervised model, called DualFair, that can debias sensitive attributes like gender and race from learned representations. Unlike existing models that target a single type of fairness, our model jointly optimizes for two fairness criteria—group fairness and counterfactual fairness—and hence makes fairer predictions at both the group and individual lev
The moisture inside the IC packages induces the several deformation failures, such as popcorn crack and swelling during the solder reflowing process. In semiconductor industry, over the past few years, the equivalent acceleration time for JEDEC moisture sensitivity level has been updated based on the weight gain measurements when the package structure and materials were modified. It costs long test times which may induce the significant delay of new product development and reliability evaluation
Anomaly detection aims at identifying deviant instances from the normal data distribution. Many advances have been made in the field, including the innovative use of unsupervised contrastive learning. However, existing methods generally assume clean training data and are limited when the data contain unknown anomalies. This paper presents Elsa, a novel semi-supervised anomaly detection approach that unifies the concept of energy-based models with unsupervised contrastive learning. Elsa instills
Satellite imagery has long been an attractive data source that provides a wealth of information on human-inhabited areas. While super resolution satellite images are rapidly becoming available, little study has focused on how to extract meaningful information about human habitation patterns and economic scales from such data. We present READ, a new approach for obtaining essential spatial representation for any given district from high-resolution satellite imagery based on deep neural networks.
This paper presents FedX, an unsupervised federated learning framework. Our model learns unbiased representation from decentralized and heterogeneous local data. It employs a two-sided knowledge distillation with contrastive learning as a core component, allowing the federated system to function without requiring clients to share any data features. Furthermore, its adaptable architecture can be used as an add-on module for existing unsupervised algorithms in federated settings. Experiments show
Large Language Models (LLMs), with their remarkable ability to tackle challenging and unseen reasoning problems, hold immense potential for tabular learning, that is vital for many real-world applications. In this paper, we propose a novel in-context learning framework, FeatLLM, which employs LLMs as feature engineers to produce an input data set that is optimally suited for tabular predictions. The generated features are used to infer class likelihood with a simple downstream machine learning m
The key technique for wireless TVs is to deliver multimedia streams stably over an unstable wireless link. In wireless LANs, a link level retransmission mechanism is used to deal with packet losses due to unstable wireless links. While the retransmission mechanism is effective to hide link level losses to users, if may causes additional delay to other packets which is fatal for real time multimedia traffic such as TV streams. In this paper, we propose a retransmission scheduling scheme to remove
Open papers in the app to read, cite, and organize with AI.