Seungeun Oh
연세대학교 첨단컴퓨팅학부 · 컴퓨터과학
Seungeun Oh 교수의 연구실은 분산학습 및 에지 컴퓨팅 환경에서의 통신 효율성과 개인정보 보호를 동시에 확보하는 혁신적인 머신러닝 프레임워크를 주요 연구 분야로 다룹니다. 특히 스플릿 러닝(Split Learning), 혼합형 플러드러닝(Mix2FLD), 하이브리드 언어모델(HLM) 등 다양한 분산 학습 아키텍처를 통해 기기 측의 자원 부담을 줄이고, 엣지 기반의 저지연·고성능 학습을 실현하고자 합니다. 또한, 업로드/다운로드 대역폭 비대칭성과 프라이버시 보호 문제를 동시에 해결하는 통신 효율적인 알고리즘 설계에 초점을 맞추고 있습니다.
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
On-device machine learning (ML) enables the training process to exploit a massive amount of user-generated private data samples. To enjoy this benefit, inter-device communication overhead should be minimized. With this end, we propose federated distillation (FD), a distributed model training algorithm whose communication payload size is much smaller than a benchmark scheme, federated learning (FL), particularly when the model size is large. Moreover, user-generated data samples are likely to bec
Cell mass and chemical composition are important aggregate cellular properties that are especially relevant to physiological processes, such as growth control and tissue homeostasis. Despite their importance, it has been difficult to measure these features quantitatively at the individual cell level in intact tissue. Here, we introduce normalized Raman imaging (NoRI), a stimulated Raman scattering (SRS) microscopy method that provides the local concentrations of protein, lipid, and water from li
Split learning (SL) is a promising distributed learning framework that enables to utilize the huge data and parallel computing resources of mobile devices. SL is built upon a model-split architecture, wherein a server stores an upper model segment that is shared by different mobile clients storing its lower model segments. Without exchanging raw data, SL achieves high accuracy and fast convergence by only uploading smashed data from clients and downloading global gradients from the server. Nonet
This letter proposes a novel communication-efficient and privacy-preserving distributed machine learning framework, coined Mix2FLD. To address uplink-downlink capacity asymmetry, local model outputs are uploaded to a server in the uplink as in federated distillation (FD), whereas global model parameters are downloaded in the downlink as in federated learning (FL). This requires a model output-to-parameter conversion at the server, after collecting additional data samples from devices. To preserv
Devices at the edge of wireless networks are the last mile data sources for machine learning (ML). As opposed to traditional ready-made public datasets, these user-generated private datasets reflect the freshest local environments in real time. They are thus indispensable for enabling mission-critical intelligent systems, ranging from fog radio access networks (RANs) to driverless cars and e-Health wearables. This article focuses on how to distill high-quality on-device ML models using fog compu
Nanogap electrodes-based dielectric spectroscopy is introduced to create ultrasensitive biomolecular sensors by minimizing the effects of electrode polarization. The electrode polarization is a major source of error in determining the impedance of biological samples in solution. The unwanted double layer impedance due to the electrode polarization impedance. is caused by the accumulation of ions on the surface of electrode. This effect becomes more dominant in low frequency region (< 1 kHz). In
Abstract Cell mass and its chemical composition are important aggregate cellular variables for physiological processes including growth control and tissue homeostasis. Despite their central importance, it has been difficult to quantitatively measure these quantities from single cells in intact tissue. Here, we introduce Normalized Raman Imaging (NoRI), a Stimulated Raman Scattering (SRS) microscopy method that provides the local concentrations of protein, lipid and water from live or fixed tissu
On-device machine learning (ML) has brought about the accessibility to a tremendous amount of data from the users while keeping their local data private instead of storing it in a central entity. However, for privacy guarantee, it is inevitable at each device to compensate for the quality of data or learning performance, especially when it has a non-IID training dataset. In this paper, we propose a data augmentation framework using a generative model: multi-hop federated augmentation with sample
To cope with the lack of on-device machine learning samples, this article presents a distributed data augmentation algorithm, coined federated data augmentation (FAug). In FAug, devices share a tiny fraction of their local data, i.e., seed samples, and collectively train a synthetic sample generator that can augment the local datasets of devices. To further improve FAug, we introduce a multihop-based seed sample collection method and an oversampling technique that mixes up collected seed samples
There have been signs of a changing perspective regarding what governments are expected to do and how they should do it since 1980s. Because the public sector is increasingly seen as rigid and bureaucratic, expensive, and inefficient, local governments in many countries are interested in developing performance evaluation system and NPM(New Public Management). This study tries to compare the local performance evaluation systems among several countries like USA, UK, Australia, Netherlands, Finland
In recent years, split learning (SL) has emerged as a promising distributed learning framework that can utilize big data in parallel without privacy leakage while reducing client-side computing resources. In the initial implementation of SL, however, the server serves multiple clients sequentially incurring high latency. Parallel implementation of SL can alleviate this latency problem, but existing Parallel SL algorithms compromise scalability due to its fundamental structural problem. To this e
This paper studies a hybrid language model (HLM) architecture that integrates a small language model (SLM) operating on a mobile device with a large language model (LLM) hosted at the base station (BS) of a wireless network. The HLM token generation process follows the speculative inference principle: the SLM’s vocabulary distribution is uploaded to the LLM, which either accepts or rejects it, with rejected tokens being resampled by the LLM. While this approach ensures alignment between the voca
Recently, vision transformer (ViT) has started to outpace the conventional CNN in computer vision tasks. Considering privacy-preserving distributed learning with ViT, federated learning (FL) communicates models, which becomes ill-suited due to ViT' s large model size and computing costs. Split learning (SL) detours this by communicating smashed data at a cut-layer, yet suffers from data privacy leakage and large communication costs caused by high similarity between ViT' s smashed data and input
Public corporations, one of the traditional governance mechanisms designed to deliver public service by central and local government, have played a significant role in the market. Therefore, the study of public corporations as a sub-field of public administration is essential to cultivation an understanding of the context of administrative reform and government performance. Despite the strong historical and institutional impact of Japanese public corporations on its system in Korea, such entitie