Skip to main content

Minsuk Jeon

Pohang University of Science and Technology · 情報科学

研究室紹介

Professor Minsuk Jeon's research lab focuses on system-level optimization for modern datacenter workloads, with a strong emphasis on improving resource efficiency, reliability, and energy sustainability. The lab investigates challenges in GPU cluster scheduling for deep learning workloads, SSD reliability in production environments, and memory energy optimization through intelligent virtual machine scheduling. Their work bridges systems research with real-world deployment, addressing critical bottlenecks in scalability, performance, and system resilience. Key contributions include novel scheduling algorithms, failure characterization frameworks, and energy-aware system design for cloud and datacenter infrastructures.

GPU schedulingSSD reliabilityenergy-efficient computingdeep learning systemsvirtual machine scheduling

Research Overview

Papers
54
Total Citations
1,108
Papers (5y)
20
Primary Field
情報科学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
20total
2021
2022
2023
2024
2026
Citations per year (5y)
88total
20212022202320242026

Selected Papers

15
1
Article|169 citations·2019
Tiresias: A GPU Cluster Manager for Distributed Deep Learning
Juncheng Gu, Mosharaf Chowdhury, Kang G. Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Harry Liu, Chuanxiong Guo
Scholarworks@UNIST (Ulsan National Institute of Science and Technology)

Deep learning (DL) training jobs bring some unique challenges to existing cluster managers, such as unpredictable training times, an all-or-nothing execution model, and inflexibility in GPU sharing. Our analysis of a large GPU cluster in production shows that existing big data schedulers cause long queueing delays and low overall performance.\nWe present Tiresias, a GPU cluster manager tailored for distributed DL training jobs, which efficiently schedules and places DL jobs to reduce their job c

Computer Networks and CommunicationsComputer Science
2
Article|142 citations·2010
Replicated abstract data types: Building blocks for collaborative applications
Hyun-Gul Roh, Myeongjae Jeon, Jin‐Soo Kim, Joonwon Lee
SJR Q1Journal of Parallel and Distributed Computing
Computer Networks and CommunicationsComputer Science
3
Preprint|118 citations·2019
Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao, Fan Yang
arXiv (Cornell University)OA

With widespread advances in machine learning, a number of large enterprises\nare beginning to incorporate machine learning models across a number of\nproducts. These models are typically trained on shared, multi-tenant GPU\nclusters. Similar to existing cluster computing workloads, scheduling\nframeworks aim to provide features like high efficiency, resource isolation,\nfair sharing across users, etc. However Deep Neural Network (DNN) based\nworkloads, predominantly trained on GPUs, differ in tw

Information SystemsComputer Science
4
Article|88 citations·2014
Predictive parallelization
Myeongjae Jeon, Saehoon Kim, Seung-won Hwang, Yuxiong He, Sameh Elnikety, Alan L. Cox, Scott Rixner

Web search engines are optimized to reduce the high-percentile response time to consistently provide fast responses to almost all user queries. This is a challenging task because the query workload exhibits large variability, consisting of many short-running queries and a few long-running queries that significantly impact the high-percentile response time. With modern multicore servers, parallelizing the processing of an individual query is a promising solution to reduce query execution time, bu

Computer Networks and CommunicationsComputer Science
5
Article|66 citations·2016
SSD Failures in Datacenters
Iyswarya Narayanan, Di Wang, Myeongjae Jeon, Bikash Sharma, Laura Caulfield, Anand Sivasubramaniam, Ben Cutler, Jie Liu, Badriddine Khessib, Kushagra Vaid
OA

Despite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, and such information cannot really help understand how SSD failures can manifest and impact the operation of production systems, in order to take appropriate remedial measures. Besides actual failure data and the symptoms exhibited by SSDs before failing, a detailed characterization effort requires wide

Computer Networks and CommunicationsComputer Science
6
Article|57 citations·2010
Energy Reduction in Consolidated Servers through Memory-Aware Virtual Machine Scheduling
Jae-Wan Jang, Myeongjae Jeon, Hyo-Sil Kim, Heeseung Jo, Jin‐Soo Kim, Seungryoul Maeng
SJR Q1IEEE Transactions on Computers

Increasing energy consumption in server consolidation environments leads to high maintenance costs for data centers. Main memory, no less than processor, is a major energy consumer in this environment. This paper proposes a technique for reducing memory energy consumption using virtual machine scheduling in multicore systems. We devise several heuristic scheduling algorithms by using a memory power simulator, which we designed and implemented. We also implement the biggest cover set first (BCSF)

Information SystemsComputer Science
7
Article|53 citations·2016
SSD Failures in Datacenters
Iyswarya Narayanan, Di Wang, Myeongjae Jeon, Bikash Sharma, Laura Caulfield, Anand Sivasubramaniam, Ben Cutler, Jie Liu, Badriddine Khessib, Kushagra Vaid

Despite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, which cannot really help understand how SSD failures can manifest and impact production systems, in order to take appropriate actions. Besides failure data, a detailed characterization requires wide spectrum of data about factors influencing SSD failures, right from provisioning (what models' where and

Computer Networks and CommunicationsComputer Science
8
Article|49 citations·2013
Adaptive parallelism for web search
Myeongjae Jeon, Yuxiong He, Sameh Elnikety, Alan L. Cox, Scott Rixner

A web search query made to Microsoft Bing is currently parallelized by distributing the query processing across many servers. Within each of these servers, the query is, however, processed sequentially. Although each server may be processing multiple queries concurrently, with modern multicore servers, parallelizing the processing of an individual query within the server may nonetheless improve the user's experience by reducing the response time. In this paper, we describe the issues that make t

Information SystemsComputer Science
9
Article|48 citations·2017
Streambox: Modern Stream Processing on a Multicore Machine
Felix Xiaozhu Lin, Gennady Pekhimenko, Heejin Park, Hongyi Xin, Kathryn S. McKinley, Myeongjae Jeon
Computer Vision and Pattern RecognitionComputer Science
10
Article|43 citations·2018
Multi-tenant GPU Clusters for Deep Learning Workloads: Analysis and Implications
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao, Fan Yang

With widespread advances in machine learning, a number of large enterprises are beginning to incorporate machine learning models across a number of products. These models are typically trained on shared, multi-tenant GPU clusters. Similar to existing cluster computing workloads, scheduling frameworks aim to provide features like high efficiency, resource isolation, fair sharing across users, etc. However Deep Neural Network (DNN) based workloads, predominantly trained on GPUs, differ in two sign

Information SystemsComputer Science
11
Book Chapter|42 citations·2008
Guest-Aware Priority-Based Virtual Machine Scheduling for Highly Consolidated Server
Dong-Sung Kim, Hwanju Kim, Myeongjae Jeon, Euiseong Seo, Joonwon Lee
SJR Q2Lecture notes in computer science
Information SystemsComputer Science
12
Article|26 citations·2016
TPC
Myeongjae Jeon, Yuxiong He, Hwanju Kim, Sameh Elnikety, Scott Rixner, Alan L. Cox

In interactive services such as web search, recommendations, games and finance, reducing the tail latency is crucial to provide fast response to every user. Using web search as a driving example, we systematically characterize interactive workload to identify the opportunities and challenges for reducing tail latency. We find that the workload consists of mainly short requests that do not benefit from parallelism, and a few long requests which significantly impact the tail but exhibit high paral

Information SystemsComputer Science
13
Article|13 citations·2012
Workload Characterization and Performance Implications of Large-Scale Blog Servers
Myeongjae Jeon, Youngjae Kim, Jeaho Hwang, Joonwon Lee, Euiseong Seo
SJR Q2ACM Transactions on the Web

With the ever-increasing popularity of Social Network Services (SNSs), an understanding of the characteristics of these services and their effects on the behavior of their host servers is critical. However, there has been a lack of research on the workload characterization of servers running SNS applications such as blog services. To fill this void, we empirically characterized real-world Web server logs collected from one of the largest South Korean blog hosting sites for 12 consecutive days. T

Computer Networks and CommunicationsComputer Science
14
Article|9 citations·2003
Fair scheduling algorithm for wireless packet networks
Myeongjae Jeon, Hiroyuki Morikawa, Tadayoshi Aoyama

In wireless packet networks, fair scheduling algorithms originally devised for wireline networks should be adapted to deal with bursty and location-dependent wireless channel errors. We present WGPS (Wireless General Processor Sharing) as a wireless fair scheduling and PWGPS (Packetized Wireless General Processor Sharing) as a packet scheduling algorithm realizing WGPS. WGPS is an extension of GPS (Generalized Processor Sharing), the fair scheduling in wired networks, and operates differently fr

Computer Networks and CommunicationsComputer Science
15
Article|7 citations·2013
Reducing DRAM row activations with eager read/write clustering
Myeongjae Jeon, Conglong Li, Alan L. Cox, Scott Rixner
SJR Q2ACM Transactions on Architecture and Code Optimization

This article describes and evaluates a new approach to optimizing DRAM performance and energy consumption that is based on eagerly writing dirty cache lines to DRAM. Under this approach, many dirty cache lines are written to DRAM before they are evicted. In particular, dirty cache lines that have not been recently accessed are eagerly written to DRAM when the corresponding row has been activated by an ordinary, noneager access, such as a read. This approach enables clustering of reads and writes

Hardware and ArchitectureComputer Science

Research Areas

Information SystemsComputer Networks and CommunicationsComputer Vision and Pattern RecognitionArtificial IntelligenceHardware and ArchitectureElectrical and Electronic Engineering

Minsuk Jeonの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。