Skip to main content
QUICK REVIEW

[論文レビュー] Artificial Intelligence for IT Operations (AIOPS) Workshop White Paper

Jasmin Bogatinovski, Sasho Nedelkoski|arXiv (Cornell University)|Jan 15, 2021
Software System Performance and Reliability参考文献 28被引用数 9
ひとこと要約

このホワイトペーパーは、インフラストラクチャ運用における人工知能(AIOPS)の現状を要約し、特に異常検出、原因分析、障害予測を含む障害管理を主な研究分野として特定している。統一された分類法を提案し、因果グラフやフェデレーテッドラーニングといった主要な技術を強調するとともに、研究の進展を促進するための公開ベンチマークの構築を提言している。

ABSTRACT

Artificial Intelligence for IT Operations (AIOps) is an emerging interdisciplinary field arising in the intersection between the research areas of machine learning, big data, streaming analytics, and the management of IT operations. AIOps, as a field, is a candidate to produce the future standard for IT operation management. To that end, AIOps has several challenges. First, it needs to combine separate research branches from other research fields like software reliability engineering. Second, novel modelling techniques are needed to understand the dynamics of different systems. Furthermore, it requires to lay out the basis for assessing: time horizons and uncertainty for imminent SLA violations, the early detection of emerging problems, autonomous remediation, decision making, support of various optimization objectives. Moreover, a good understanding and interpretability of these aiding models are important for building trust between the employed tools and the domain experts. Finally, all this will result in faster adoption of AIOps, further increase the interest in this research field and contribute to bridging the gap towards fully-autonomous operating IT systems. The main aim of the AIOPS workshop is to bring together researchers from both academia and industry to present their experiences, results, and work in progress in this field. The workshop aims to strengthen the community and unite it towards the goal of joining the efforts for solving the main challenges the field is currently facing. A consensus and adoption of the principles of openness and reproducibility will boost the research in this emerging area significantly.

研究の動機と目的

  • 機械学習、ビッグデータ、IT運用を統合する分野として、急速に発展する多様な分野であるAIOPSを統合的かつ体系的に整理すること。
  • AIOPSにおける主な課題、例えばモデルの解釈可能性、不確実性の評価、自律的修復の実現を特定すること。
  • 研究の革新を加速させるために、コミュニティ全体でオープンかつ再現可能な研究手法を採用することを促進すること。
  • AIOPS手法の相互比較を可能にし、研究の進捗を追跡できる共通のベンチマークフレームワークを確立すること。
  • より良い原因分析と自己修復システムを実現することで、完全に自律的なIT運用への道筋を前進させること。

提案手法

  • 1,000件を超えるAIOPS関連論文を系統的に調査し、研究分野をマクロ分野、カテゴリ、時間的トレンドごとに分類した。
  • AIOPS応用分野の分類法を提案し、その中で障害管理が主要なカテゴリであり、さらに障害検出、予測、原因分析、修復に細分化されていることを示した。
  • ハーケス過程やグラフノード埋め込みといった因果モデル化技術を用いて、アラームの因果関係を推定し、原因アラームを特定した。
  • ディープラーニングとサービス依存性グラフを応用し、クラウドマイクロサービスにおけるパフォーマンス劣化要因を高い精度(0.92)で同定した。
  • 教師-生徒蒸留を用いた分散型フェデレーテッドラーニング手法を開発し、生のログデータやモデルパラメータを露呈せずに異常検出モデルを共有した。
  • 人工スウォームインテリジェンスフレームワークを導入し、パブリッククラウドにおけるリソース共有を最適化することで、QoE(品質保証)を維持しながらリソース利用効率を向上させた。

実験結果

リサーチクエスチョン

  • RQ1AIOPS分野における主な応用分野と研究トレンドは、時間的経過と分野ごとにどのように変化しているか?
  • RQ2因果推論とグラフベースモデリングは、複雑なITシステムにおける原因分析をどのように改善できるか?
  • RQ3信頼性があり、解釈可能で、自律的なAI駆動型IT運用を実現する上で、どのような主な課題が存在するか?
  • RQ4機密データを露呈せずに、ITサービス間でプライバシーを守った知識共有をどのように実現できるか?
  • RQ5公開ベンチマークは、AIOPS研究における公平な比較と研究進捗の加速にどのような役割を果たせるか?

主な発見

  • 障害管理は、AIOPS関連論文全体の62.1%を占めており、その中でも障害検出(33.7%)と原因分析(26.7%)が最も活発な研究分野である。
  • 近年、障害検出分野の研究が急増しており、2018年から2019年の間に71件の論文が発表されたが、これはリソース割り当て分野全体(69件)を上回っている。
  • サービス依存性グラフとメトリクスの異常を組み合わせたディープラーニングベースのマイクロサービスパフォーマンス診断手法は、原因特定において0.92の精度を達成した。
  • フェデレーテッドラーニング手法により、トレーニングデータやモデルパラメータを共有することなく、ログベースの異常検出が向上し、プライバシーが保たれた。
  • 人工スウォームインテリジェンスモデルは、ピーク負荷時においてもQoEを維持しながら、複数の顧客間でクラウドリソースの割り当てを効果的にバランスさせた。
  • 関心の高まりにもかかわらず、AIOPS手法の相互比較を可能にする公開ベンチマークは存在せず、研究の進捗と再現性の面で障害要因となっている。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。