전명재 교수
Minsuk Jeon
포항공과대학교 컴퓨터공학과 · 컴퓨터과학
연구실 소개
전명재 교수의 연구실은 데이터센터 환경에서의 자원 관리 및 시스템 신뢰성에 중점을 두고 있으며, 특히 딥러닝 워크로드를 위한 고성능 GPU 클러스터 스케줄링 기술과 SSD의 실재 환경에서의 신뢰성 특성 분석을 핵심 연구 분야로 다룹니다. 대용량 데이터 처리 환경에서의 에너지 효율성 향상 및 가상머신 스케줄링 기법을 통해 서버 통합 환경의 전력 소비를 최적화하는 연구도 진행 중입니다. 특히 실사용 환경에서의 장비 고장 원인과 영향을 체계적으로 분석함으로써, 생산성과 안정성을 동시에 높이는 시스템 설계 기반의 솔루션 개발을 추구합니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15Deep learning (DL) training jobs bring some unique challenges to existing cluster managers, such as unpredictable training times, an all-or-nothing execution model, and inflexibility in GPU sharing. Our analysis of a large GPU cluster in production shows that existing big data schedulers cause long queueing delays and low overall performance.\nWe present Tiresias, a GPU cluster manager tailored for distributed DL training jobs, which efficiently schedules and places DL jobs to reduce their job c
With widespread advances in machine learning, a number of large enterprises\nare beginning to incorporate machine learning models across a number of\nproducts. These models are typically trained on shared, multi-tenant GPU\nclusters. Similar to existing cluster computing workloads, scheduling\nframeworks aim to provide features like high efficiency, resource isolation,\nfair sharing across users, etc. However Deep Neural Network (DNN) based\nworkloads, predominantly trained on GPUs, differ in tw
Web search engines are optimized to reduce the high-percentile response time to consistently provide fast responses to almost all user queries. This is a challenging task because the query workload exhibits large variability, consisting of many short-running queries and a few long-running queries that significantly impact the high-percentile response time. With modern multicore servers, parallelizing the processing of an individual query is a promising solution to reduce query execution time, bu
Despite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, and such information cannot really help understand how SSD failures can manifest and impact the operation of production systems, in order to take appropriate remedial measures. Besides actual failure data and the symptoms exhibited by SSDs before failing, a detailed characterization effort requires wide
Increasing energy consumption in server consolidation environments leads to high maintenance costs for data centers. Main memory, no less than processor, is a major energy consumer in this environment. This paper proposes a technique for reducing memory energy consumption using virtual machine scheduling in multicore systems. We devise several heuristic scheduling algorithms by using a memory power simulator, which we designed and implemented. We also implement the biggest cover set first (BCSF)
Despite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, which cannot really help understand how SSD failures can manifest and impact production systems, in order to take appropriate actions. Besides failure data, a detailed characterization requires wide spectrum of data about factors influencing SSD failures, right from provisioning (what models' where and
A web search query made to Microsoft Bing is currently parallelized by distributing the query processing across many servers. Within each of these servers, the query is, however, processed sequentially. Although each server may be processing multiple queries concurrently, with modern multicore servers, parallelizing the processing of an individual query within the server may nonetheless improve the user's experience by reducing the response time. In this paper, we describe the issues that make t
With widespread advances in machine learning, a number of large enterprises are beginning to incorporate machine learning models across a number of products. These models are typically trained on shared, multi-tenant GPU clusters. Similar to existing cluster computing workloads, scheduling frameworks aim to provide features like high efficiency, resource isolation, fair sharing across users, etc. However Deep Neural Network (DNN) based workloads, predominantly trained on GPUs, differ in two sign
In interactive services such as web search, recommendations, games and finance, reducing the tail latency is crucial to provide fast response to every user. Using web search as a driving example, we systematically characterize interactive workload to identify the opportunities and challenges for reducing tail latency. We find that the workload consists of mainly short requests that do not benefit from parallelism, and a few long requests which significantly impact the tail but exhibit high paral
With the ever-increasing popularity of Social Network Services (SNSs), an understanding of the characteristics of these services and their effects on the behavior of their host servers is critical. However, there has been a lack of research on the workload characterization of servers running SNS applications such as blog services. To fill this void, we empirically characterized real-world Web server logs collected from one of the largest South Korean blog hosting sites for 12 consecutive days. T
In wireless packet networks, fair scheduling algorithms originally devised for wireline networks should be adapted to deal with bursty and location-dependent wireless channel errors. We present WGPS (Wireless General Processor Sharing) as a wireless fair scheduling and PWGPS (Packetized Wireless General Processor Sharing) as a packet scheduling algorithm realizing WGPS. WGPS is an extension of GPS (Generalized Processor Sharing), the fair scheduling in wired networks, and operates differently fr
This article describes and evaluates a new approach to optimizing DRAM performance and energy consumption that is based on eagerly writing dirty cache lines to DRAM. Under this approach, many dirty cache lines are written to DRAM before they are evicted. In particular, dirty cache lines that have not been recently accessed are eagerly written to DRAM when the corresponding row has been activated by an ordinary, noneager access, such as a read. This approach enables clustering of reads and writes
대표 연구 분야
전명재 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.