Skip to main content
QUICK REVIEW

[논문 리뷰] Sieve: Actionable Insights from Monitored Metrics in Microservices

Jörg Thalheim, António Rodrigues|arXiv (Cornell University)|2017. 09. 20.
Software System Performance and Reliability참고 문헌 48인용 수 11
한 줄 요약

Sieve는 미세 서비스에서 모니터링된 메트릭의 차원을 줄이고 Granger 인과성(Granger Causality)을 사용하여 인과적 의존성을 추론함으로써 실행 가능한 인사이트를 도출하는 플랫폼이다. 이는 메트릭을 10~100배 줄이고, 스토리지에서 최대 90%, CPU에서 최대 80%의 시스템 오버헤드를 감소시키며, 인스트루멘테이션 또는 사전 지식 없이도 효과적인 오토스케일링과 장애 원인 분석을 가능하게 한다.

ABSTRACT

Major cloud computing operators provide powerful monitoring tools to understand the current (and prior) state of the distributed systems deployed in their infrastructure. While such tools provide a detailed monitoring mechanism at scale, they also pose a significant challenge for the application developers/operators to transform the huge space of monitored metrics into useful insights. These insights are essential to build effective management tools for improving the efficiency, resiliency, and dependability of distributed systems. This paper reports on our experience with building and deploying Sieve - a platform to derive actionable insights from monitored metrics in distributed systems. Sieve builds on two core components: a metrics reduction framework, and a metrics dependency extractor. More specifically, Sieve first reduces the dimensionality of metrics by automatically filtering out unimportant metrics by observing their signal over time. Afterwards, Sieve infers metrics dependencies between distributed components of the system using a predictive-causality model by testing for Granger Causality. We implemented Sieve as a generic platform and deployed it for two microservices-based distributed systems: OpenStack and ShareLatex. Our experience shows that (1) Sieve can reduce the number of metrics by at least an order of magnitude (10 - 100$ imes$), while preserving the statistical equivalence to the total number of monitored metrics; (2) Sieve can dramatically improve existing monitoring infrastructures by reducing the associated overheads over the entire system stack (CPU - 80%, storage - 90%, and network - 50%); (3) Lastly, Sieve can be effective to support a wide-range of workflows in distributed systems - we showcase two such workflows: orchestration of autoscaling, and Root Cause Analysis (RCA).

연구 동기 및 목표

  • 고차원적이고 모니터링된 메트릭을 미세 서비스 관리에 활용 가능한 인사이트로 전환하는 데 도전하는 것.
  • 통계적 관련성을 유지하면서 모니터링되는 메트릭 수를 줄여 모니터링 오버헤드를 최소화하는 것.
  • 분산된 구성 요소 간 메트릭 간의 인과 관계를 추론하여 시스템 관리 워크플로우를 지원하는 것.
  • 기존 모니터링 데이터에서 비정규화된, 일반적이고 애플리케이션에 관계없이 적용 가능한 인사이트 유도를 가능하게 하는 것.
  • 오토스케일링 및 장애 원인 분석과 같은 핵심 DevOps 워크플로우를 추론된 메트릭 의존성으로 지원하는 것.

제안 방법

  • Sieve는 메트릭의 시간적 신호 강도와 분산을 분석하여 중요하지 않은 메트릭을 필터링하는 메트릭 감소 엔진을 사용한다.
  • Granger 인과성 기반의 예측-인과 모델을 적용하여 서로 다른 구성 요소의 메트릭 간 인과 관계를 추론한다.
  • 사전에 메트릭 시계열이나 애플리케이션 의미에 대한 지식이 필요 없이 비지도 학습 방식으로 작동한다.
  • 실제 미세 서비스 환경에서 추론된 인과 모델의 타당성을 검증하기 위해 워크로드 생성기나 프로덕션 트레이스를 사용한다.
  • OpenStack과 ShareLatex에 배포되어 다양한 시스템에서의 확장성과 적응 가능성에 대한 증명을 한다.
  • 모니터링 대상 시스템에 대한 인스트루멘테이션이나 변경 없이 기존 모니터링 스택과 통합된다.

실험 결과

연구 질문

  • RQ1통계적 관련성을 유지하면서 미세 서비스의 모니터링 메트릭 차원을 자동으로 줄일 수 있는가?
  • RQ2시계열 데이터만을 사용하여 분산된 구성 요소 간 메트릭 간 의미 있는 인과적 의존성을 추론할 수 있는가?
  • RQ3메트릭 감소와 의존성 추론이 실제 시스템에서 모니터링 오버헤드를 어느 정도 감소시킬 수 있는가?
  • RQ4추론된 메트릭 의존성이 오토스케일링 및 장애 원인 분석과 같은 실제 관리 워크플로우를 효과적으로 지원할 수 있는가?
  • RQ5이러한 시스템은 다양한 미세 서비스 아키텍처에서 일반적이고, 비지도 학습 방식이며 침습적이지 않게 어떻게 구현할 수 있는가?

주요 결과

  • Sieve는 전체 메트릭 세트와 통계적 동등성을 유지하면서도 모니터링되는 메트릭 수를 최소 10배에서 최대 100배까지 줄인다.
  • 시스템 전체 스택에서 CPU 오버헤드는 최대 80%, 스토리지 오버헤드는 최대 90%, 네트워크 사용량은 50% 감소한다.
  • 추론된 인과 모델은 자원 수요를 이끄는 핵심 메트릭을 식별함으로써 효과적인 오토스케일링을 가능하게 한다.
  • 플랫폼은 장애가 발생하기 시작한 최초의 메트릭을 식별함으로써 빠른 장애 진단을 지원하는 장애 원인 분석을 가능하게 한다.
  • OpenStack과 ShareLatex에서의 사례를 통해 Sieve의 접근 방식은 다양한 미세 서비스 시스템에서 일반적이고 재사용 가능하다.
  • 시스템은 인스트루멘테이션 또는 사전 지식 없이 관측된 메트릭 시계열 데이터에만 의존하여 작동한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.