Skip to main content
QUICK REVIEW

[論文レビュー] Performance-Aware Management of Cloud Resources: A Taxonomy and Future Directions

Sara Kardani Moghaddam, Rajkumar Buyya|arXiv (Cornell University)|Aug 7, 2018
Software System Performance and Reliability参考文献 65被引用数 9
ひとこと要約

本稿は、パフォーマンスに配慮したクラウドリソース管理の包括的な分類体系と今後の研究方向性を提案する。データ分析(異常検出、ワークロード予測)と動的リソース調整(自動スケーリング)を統合し、リアルタイム監視、適応的設定、アプリケーション固有の検出精度の課題を特定する。動的なクラウド環境においてSLAを維持するには、エンドツーエンドでデータドリブンかつ適応的であることが不可欠であることを強調する。

ABSTRACT

Dynamic nature of the cloud environment has made distributed resource management process a challenge for cloud service providers. The importance of maintaining the quality of service in accordance with customer expectations as well as the highly dynamic nature of cloud-hosted applications add new levels of complexity to the process. Advances to the big data learning approaches have shifted conventional static capacity planning solutions to complex performance-aware resource management methods. It is shown that the process of decision making for resource adjustment is closely related to the behaviour of the system including the utilization of resources and application components. Therefore, a continuous monitoring of system attributes and performance metrics provide the raw data for the analysis of problems affecting the performance of the application. Data analytic methods such as statistical and machine learning approaches offer the required concepts, models and tools to dig into the data, find general rules, patterns and characteristics that define the functionality of the system. Obtained knowledge form the data analysis process helps to find out about the changes in the workloads, faulty components or problems that can cause system performance to degrade. A timely reaction to performance degradations can avoid violations of the service level agreements by performing proper corrective actions including auto-scaling or other resource adjustment solutions. In this paper, we investigate the main requirements and limitations in cloud resource management including a study of the approaches in workload and anomaly analysis in the context of the performance management in the cloud. A taxonomy of the works on this problem is presented which identifies the main approaches in existing researches from data analysis side to resource adjustment techniques.

研究の動機と目的

  • サービスレベル契約(SLA)の遵守を維持しながら、動的なクラウドワークロードを管理する複雑さの増大に対処すること。
  • 動的で多様なクラウドワークロード下で、従来の静的およびヒューリスティックベースのリソース管理手法に見られる限界を特定すること。
  • データ分析(異常検出、ワークロード予測)と自動リソース調整(自動スケーリング)を統合し、包括的なパフォーマンス管理を実現すること。
  • リアルタイム感度、適応的設定、アプリケーション固有の検出精度のトレードオフにおける研究ギャップを浮き彫りにすること。
  • 今後の研究の指針となるよう、データ収集、分析、リソース調整をカバーする構造的分類体系を提供すること。

提案手法

  • アーキテクチャ、データ粒度、パフォーマンス問題の種別、リソース管理行動に基づいてアプローチを分類する多次元分類体系を提案する。
  • ワークロード分析、異常検出、自動スケーリング分野の既存研究を調査し、データ分析とリソース管理の統合に焦点を当てる。
  • データドリブン意思決定パイプライン(監視 → データ収集 → 分析(統計および機械学習) → 矯正措置)を分析する。
  • ビッグデータ分析がシステムメトリクスからパターン、トレンド、パフォーマンス劣化の兆候を解明する役割を強調する。
  • 変化するクラウドワークロードに適応するため、動的しきい値チューニングと学習モデルの自動設定の必要性を提示する。
  • フィードバックループを介した統合的異常原因推定を提唱し、リソース管理における計画立案と行動選択の改善を図る。

実験結果

リサーチクエスチョン

  • RQ1データ分析技術を自動リソース管理と効果的に統合することで、クラウドのパフォーマンスとSLA準拠性をどのように向上させられるか?
  • RQ2実世界のクラウド環境に適用した場合、現在の異常検出法およびワークロード予測法にどのような主な限界があるか?
  • RQ3動的設定と適応的しきい値設定は、クラウド監視における機械学習モデルのパフォーマンスをどのように向上させるか?
  • RQ4リアルタイムでアプリケーションに配慮した異常検出と原因推定における、パフォーマンス管理における重要なギャップは何か?
  • RQ5不均衡なクラウドパフォーマンスデータにおいて、AUCとPRAUCといった検出精度指標を比較すると、どちらが実運用環境に適しているか?

主な発見

  • 既存の手法はしばしばデータ分析とリソース管理を別々のモジュールとして扱っており、エンドツーエンドの統合が欠けている。
  • 従来の静的設定による異常検出およびスケーリングアルゴリズムは、クラウドワークロードの動的特性に適応できない。
  • 異常原因の推定は粗い粒度であり、計画モジュールから切り離されているため、是正措置の効果が制限される。
  • 正常と異常データのインスタンス数の不均衡により、標準的な評価指標(AUC)はバイアスを受けるため、PRAUCが実世界のクラウドアプリケーションにおいてより適切な指標である。
  • 自動スケーリングシステムの現実的なパフォーマンス評価には、分散コンponent間の複雑な相互作用を考慮し、実環境でのデプロイが必要である。
  • 今後の研究では、特にディスク障害対策などの高コストな回復処理を伴う場合に、アプリケーション固有の検出精度のトレードオフを優先すべきである。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。