Minsuk Jeon
Pohang University of Science and Technology · Computer Science
About the Lab
Professor Minsuk Jeon's research lab focuses on system-level optimization for modern datacenter workloads, with a strong emphasis on improving resource efficiency, reliability, and energy sustainability. The lab investigates challenges in GPU cluster scheduling for deep learning workloads, SSD reliability in production environments, and memory energy optimization through intelligent virtual machine scheduling. Their work bridges systems research with real-world deployment, addressing critical bottlenecks in scalability, performance, and system resilience. Key contributions include novel scheduling algorithms, failure characterization frameworks, and energy-aware system design for cloud and datacenter infrastructures.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15Deep learning (DL) training jobs bring some unique challenges to existing cluster managers, such as unpredictable training times, an all-or-nothing execution model, and inflexibility in GPU sharing. Our analysis of a large GPU cluster in production shows that existing big data schedulers cause long queueing delays and low overall performance.\nWe present Tiresias, a GPU cluster manager tailored for distributed DL training jobs, which efficiently schedules and places DL jobs to reduce their job c
With widespread advances in machine learning, a number of large enterprises\nare beginning to incorporate machine learning models across a number of\nproducts. These models are typically trained on shared, multi-tenant GPU\nclusters. Similar to existing cluster computing workloads, scheduling\nframeworks aim to provide features like high efficiency, resource isolation,\nfair sharing across users, etc. However Deep Neural Network (DNN) based\nworkloads, predominantly trained on GPUs, differ in tw
Web search engines are optimized to reduce the high-percentile response time to consistently provide fast responses to almost all user queries. This is a challenging task because the query workload exhibits large variability, consisting of many short-running queries and a few long-running queries that significantly impact the high-percentile response time. With modern multicore servers, parallelizing the processing of an individual query is a promising solution to reduce query execution time, bu
Despite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, and such information cannot really help understand how SSD failures can manifest and impact the operation of production systems, in order to take appropriate remedial measures. Besides actual failure data and the symptoms exhibited by SSDs before failing, a detailed characterization effort requires wide
Increasing energy consumption in server consolidation environments leads to high maintenance costs for data centers. Main memory, no less than processor, is a major energy consumer in this environment. This paper proposes a technique for reducing memory energy consumption using virtual machine scheduling in multicore systems. We devise several heuristic scheduling algorithms by using a memory power simulator, which we designed and implemented. We also implement the biggest cover set first (BCSF)
Despite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, which cannot really help understand how SSD failures can manifest and impact production systems, in order to take appropriate actions. Besides failure data, a detailed characterization requires wide spectrum of data about factors influencing SSD failures, right from provisioning (what models' where and
A web search query made to Microsoft Bing is currently parallelized by distributing the query processing across many servers. Within each of these servers, the query is, however, processed sequentially. Although each server may be processing multiple queries concurrently, with modern multicore servers, parallelizing the processing of an individual query within the server may nonetheless improve the user's experience by reducing the response time. In this paper, we describe the issues that make t
With widespread advances in machine learning, a number of large enterprises are beginning to incorporate machine learning models across a number of products. These models are typically trained on shared, multi-tenant GPU clusters. Similar to existing cluster computing workloads, scheduling frameworks aim to provide features like high efficiency, resource isolation, fair sharing across users, etc. However Deep Neural Network (DNN) based workloads, predominantly trained on GPUs, differ in two sign
In interactive services such as web search, recommendations, games and finance, reducing the tail latency is crucial to provide fast response to every user. Using web search as a driving example, we systematically characterize interactive workload to identify the opportunities and challenges for reducing tail latency. We find that the workload consists of mainly short requests that do not benefit from parallelism, and a few long requests which significantly impact the tail but exhibit high paral
With the ever-increasing popularity of Social Network Services (SNSs), an understanding of the characteristics of these services and their effects on the behavior of their host servers is critical. However, there has been a lack of research on the workload characterization of servers running SNS applications such as blog services. To fill this void, we empirically characterized real-world Web server logs collected from one of the largest South Korean blog hosting sites for 12 consecutive days. T
In wireless packet networks, fair scheduling algorithms originally devised for wireline networks should be adapted to deal with bursty and location-dependent wireless channel errors. We present WGPS (Wireless General Processor Sharing) as a wireless fair scheduling and PWGPS (Packetized Wireless General Processor Sharing) as a packet scheduling algorithm realizing WGPS. WGPS is an extension of GPS (Generalized Processor Sharing), the fair scheduling in wired networks, and operates differently fr
This article describes and evaluates a new approach to optimizing DRAM performance and energy consumption that is based on eagerly writing dirty cache lines to DRAM. Under this approach, many dirty cache lines are written to DRAM before they are evicted. In particular, dirty cache lines that have not been recently accessed are eagerly written to DRAM when the corresponding row has been activated by an ordinary, noneager access, such as a read. This approach enables clustering of reads and writes
Research Areas
Dive deeper into Minsuk Jeon's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.