Skip to main content
QUICK REVIEW

[논문 리뷰] DRESS: Dynamic RESource-reservation Scheme for Congested Data-intensive Computing Platforms

Ying Mao, Victoria Green|arXiv (Cornell University)|2018. 05. 21.
Cloud Computing and Resource Management인용 수 5
한 줄 요약

DRESS는 데이터 집약적 컴퓨팅 플랫폼을 위한 동적 자원 예약 기법으로, 작업을 소규모 및 대규모 수요 유형으로 분류하고 각 유형에 전용 자원을 할당하며, 실시간 작업 수요와 예상 해제 패턴에 기반해 예약 비율을 동적으로 조정한다. 이는 Hadoop YARN 환경에서 평균 소작업 완료 시간을 최대 76.1% 감소시키면서도 전체 시스템 성능을 안정적으로 유지한다.

ABSTRACT

In the past few years, we have envisioned an increasing number of businesses start driving by big data analytics, such as Amazon recommendations and Google Advertisements. At the back-end side, the businesses are powered by big data processing platforms to quickly extract information and make decisions. Running on top of a computing cluster, those platforms utilize scheduling algorithms to allocate resources. An efficient scheduler is crucial to the system performance due to limited resources, e.g. CPU and Memory, and a large number of user demands. However, besides requests from clients and current status of the system, it has limited knowledge about execution length of the running jobs, and incoming jobs' resource demands, which make assigning resources a challenging task. If most of the resources are occupied by a long-running job, other jobs will have to keep waiting until it releases them. This paper presents a new scheduling strategy, named DRESS that particularly aims to optimize the allocation among jobs with various demands. Specifically, it classifies the jobs into two categories based on their requests, reserves a portion of resources for each of category, and dynamically adjusts the reserved ratio by monitoring the pending requests and estimating release patterns of running jobs. The results demonstrate DRESS significantly reduces the completion time for one category, up to 76.1% in our experiments, and in the meanwhile, maintains a stable overall system performance.

연구 동기 및 목표

  • 혼잡한 데이터 집약적 클러스터에서 소자원 작업의 긴 대기 시간과 열악한 성능 문제를 해결하기 위해.
  • 전체 시스템 안정성에 영향을 주지 않으면서도 소수요 작업의 작업 완료 시간을 단축하기 위해.
  • 실시간 작업 큐 동적 변화와 자원 해제 패턴에 적응하는 동적 자원 예약 메커니즘을 설계하기 위해.
  • Hadoop YARN 환경에서 병렬성과 자원 활용도를 향상시키는 스케줄러를 구현하고 평가하기 위해.

제안 방법

  • 입력 작업을 자원 수요 기반으로 두 유형(소규모 및 대규모)으로 분류한다.
  • 동적 예약 비율을 사용해 각 작업 유형에 대해 가변적인 시스템 자원 비율을 할당한다.
  • 보류 중인 작업 수요 모니터링 및 실행 중인 작업의 예상 해제 시간을 기반으로 예약 비율을 동적으로 조정한다.
  • 실행 중인 작업이 자원을 해제할 시점을 예측하기 위해 자원 해제 패턴 추정을 수행한다.
  • 알고리즘 3을 사용해 큐 내 소작업 비율에 따라 예약 비율을 제어한다.
  • MapReduce 및 Spark 워크로드를 모두 지원하는 Hadoop YARN 플랫폼에 구현하여 종단 간 평가를 수행한다.

실험 결과

연구 질문

  • RQ1고도로 혼잡한 데이터 집약적 클러스터에서 소수요 작업에 대한 자원 할당을 어떻게 개선할 수 있는가?
  • RQ2동적 자원 예약이 작업 완료 시간과 전반적인 시스템 성능에 미치는 영향은 무엇인가?
  • RQ3실시간 작업 큐 및 자원 해제 패턴에 기반해 예약 비율을 조정할 경우 스케줄링 효율성에 어떤 영향을 미치는가?
  • RQ4동적 자원 예약 전략이 소작업의 대기 시간을 얼마나 줄일 수 있으며, 이로 인해 대작업이나 전체 시스템 안정성에 악영향을 주지 않는가?

주요 결과

  • 40%의 소작업이 포함된 혼합 워크로드 환경에서 DRESS는 소작업의 평균 완료 시간을 최대 76.1% 감소시켰다.
  • Spark-on-YARN 실험에서 소작업의 평균 완료 시간은 51.2% 감소했고, 최대 76.1%까지 감소했다.
  • Hadoop YARN 실험에서 20개의 작업을 대상으로 소작업의 평균 완료 시간은 25.7% 감소했다.
  • 전체 시스템 메이크스팬은 안정적으로 유지되었으며, DRESS는 1035.2초를 기록했고, Capacity 스케줄러는 1028.6초를 기록했다.
  • 일부 대작업(예: 작업 7)은 완료 시간이 증가했으며(최대 29.3% 더 오래), 그러나 20개의 작업 중 12개는 평균적으로 18.5% 감소한 완료 시간을 기록했다.
  • 특히 작업 9, 12, 13의 성능이 크게 향상되어 각각 완료 시간이 23.2%, 17.5%, 10.0% 감소했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.