Yonsei University · Computer Science
Professor Seungeun Oh's research lab specializes in communication-efficient and privacy-preserving distributed machine learning, with a focus on enabling scalable, low-latency AI collaboration across resource-constrained devices. The lab develops innovative frameworks such as Split Learning, Mix2FLD, and hybrid language modeling architectures that balance model accuracy, data privacy, and system efficiency in wireless and mobile environments. Key research directions include optimizing uplink-downlink communication asymmetry, enhancing model scalability through novel aggregation and mixing techniques, and integrating small and large language models for edge intelligence. The lab also explores physical-layer protocols for high-mobility wireless networks, particularly in mmWave V2X communications.
Figures are computed from collected data and may differ slightly.
On-device machine learning (ML) enables the training process to exploit a massive amount of user-generated private data samples. To enjoy this benefit, inter-device communication overhead should be minimized. With this end, we propose federated distillation (FD), a distributed model training algorithm whose communication payload size is much smaller than a benchmark scheme, federated learning (FL), particularly when the model size is large. Moreover, user-generated data samples are likely to bec
Cell mass and chemical composition are important aggregate cellular properties that are especially relevant to physiological processes, such as growth control and tissue homeostasis. Despite their importance, it has been difficult to measure these features quantitatively at the individual cell level in intact tissue. Here, we introduce normalized Raman imaging (NoRI), a stimulated Raman scattering (SRS) microscopy method that provides the local concentrations of protein, lipid, and water from li
Split learning (SL) is a promising distributed learning framework that enables to utilize the huge data and parallel computing resources of mobile devices. SL is built upon a model-split architecture, wherein a server stores an upper model segment that is shared by different mobile clients storing its lower model segments. Without exchanging raw data, SL achieves high accuracy and fast convergence by only uploading smashed data from clients and downloading global gradients from the server. Nonet
This letter proposes a novel communication-efficient and privacy-preserving distributed machine learning framework, coined Mix2FLD. To address uplink-downlink capacity asymmetry, local model outputs are uploaded to a server in the uplink as in federated distillation (FD), whereas global model parameters are downloaded in the downlink as in federated learning (FL). This requires a model output-to-parameter conversion at the server, after collecting additional data samples from devices. To preserv
Devices at the edge of wireless networks are the last mile data sources for machine learning (ML). As opposed to traditional ready-made public datasets, these user-generated private datasets reflect the freshest local environments in real time. They are thus indispensable for enabling mission-critical intelligent systems, ranging from fog radio access networks (RANs) to driverless cars and e-Health wearables. This article focuses on how to distill high-quality on-device ML models using fog compu
Nanogap electrodes-based dielectric spectroscopy is introduced to create ultrasensitive biomolecular sensors by minimizing the effects of electrode polarization. The electrode polarization is a major source of error in determining the impedance of biological samples in solution. The unwanted double layer impedance due to the electrode polarization impedance. is caused by the accumulation of ions on the surface of electrode. This effect becomes more dominant in low frequency region (< 1 kHz). In
Abstract Cell mass and its chemical composition are important aggregate cellular variables for physiological processes including growth control and tissue homeostasis. Despite their central importance, it has been difficult to quantitatively measure these quantities from single cells in intact tissue. Here, we introduce Normalized Raman Imaging (NoRI), a Stimulated Raman Scattering (SRS) microscopy method that provides the local concentrations of protein, lipid and water from live or fixed tissu
On-device machine learning (ML) has brought about the accessibility to a tremendous amount of data from the users while keeping their local data private instead of storing it in a central entity. However, for privacy guarantee, it is inevitable at each device to compensate for the quality of data or learning performance, especially when it has a non-IID training dataset. In this paper, we propose a data augmentation framework using a generative model: multi-hop federated augmentation with sample
To cope with the lack of on-device machine learning samples, this article presents a distributed data augmentation algorithm, coined federated data augmentation (FAug). In FAug, devices share a tiny fraction of their local data, i.e., seed samples, and collectively train a synthetic sample generator that can augment the local datasets of devices. To further improve FAug, we introduce a multihop-based seed sample collection method and an oversampling technique that mixes up collected seed samples
There have been signs of a changing perspective regarding what governments are expected to do and how they should do it since 1980s. Because the public sector is increasingly seen as rigid and bureaucratic, expensive, and inefficient, local governments in many countries are interested in developing performance evaluation system and NPM(New Public Management). This study tries to compare the local performance evaluation systems among several countries like USA, UK, Australia, Netherlands, Finland
In recent years, split learning (SL) has emerged as a promising distributed learning framework that can utilize big data in parallel without privacy leakage while reducing client-side computing resources. In the initial implementation of SL, however, the server serves multiple clients sequentially incurring high latency. Parallel implementation of SL can alleviate this latency problem, but existing Parallel SL algorithms compromise scalability due to its fundamental structural problem. To this e
This paper studies a hybrid language model (HLM) architecture that integrates a small language model (SLM) operating on a mobile device with a large language model (LLM) hosted at the base station (BS) of a wireless network. The HLM token generation process follows the speculative inference principle: the SLM’s vocabulary distribution is uploaded to the LLM, which either accepts or rejects it, with rejected tokens being resampled by the LLM. While this approach ensures alignment between the voca
Recently, vision transformer (ViT) has started to outpace the conventional CNN in computer vision tasks. Considering privacy-preserving distributed learning with ViT, federated learning (FL) communicates models, which becomes ill-suited due to ViT' s large model size and computing costs. Split learning (SL) detours this by communicating smashed data at a cut-layer, yet suffers from data privacy leakage and large communication costs caused by high similarity between ViT' s smashed data and input
Public corporations, one of the traditional governance mechanisms designed to deliver public service by central and local government, have played a significant role in the market. Therefore, the study of public corporations as a sub-field of public administration is essential to cultivation an understanding of the context of administrative reform and government performance. Despite the strong historical and institutional impact of Japanese public corporations on its system in Korea, such entitie
Open papers in the app to read, cite, and organize with AI.