Skip to main content
QUICK REVIEW

[论文解读] Artificial Intelligence for IT Operations (AIOPS) Workshop White Paper

Jasmin Bogatinovski, Sasho Nedelkoski|arXiv (Cornell University)|Jan 15, 2021
Software System Performance and Reliability参考文献 28被引用 9
一句话总结

本白皮书综合了IT运维人工智能(AIOPS)的最新进展,将故障管理——尤其是异常检测、根本原因分析和故障预测——确定为研究的核心焦点。该白皮书提出了一套统一的分类体系,强调了因果图和联邦学习等关键技术,并呼吁建立公共基准,以加速实现自主IT运维的进程。

ABSTRACT

Artificial Intelligence for IT Operations (AIOps) is an emerging interdisciplinary field arising in the intersection between the research areas of machine learning, big data, streaming analytics, and the management of IT operations. AIOps, as a field, is a candidate to produce the future standard for IT operation management. To that end, AIOps has several challenges. First, it needs to combine separate research branches from other research fields like software reliability engineering. Second, novel modelling techniques are needed to understand the dynamics of different systems. Furthermore, it requires to lay out the basis for assessing: time horizons and uncertainty for imminent SLA violations, the early detection of emerging problems, autonomous remediation, decision making, support of various optimization objectives. Moreover, a good understanding and interpretability of these aiding models are important for building trust between the employed tools and the domain experts. Finally, all this will result in faster adoption of AIOps, further increase the interest in this research field and contribute to bridging the gap towards fully-autonomous operating IT systems. The main aim of the AIOPS workshop is to bring together researchers from both academia and industry to present their experiences, results, and work in progress in this field. The workshop aims to strengthen the community and unite it towards the goal of joining the efforts for solving the main challenges the field is currently facing. A consensus and adoption of the principles of openness and reproducibility will boost the research in this emerging area significantly.

研究动机与目标

  • 整合并构建新兴的、跨学科的AIOPS领域,该领域融合了机器学习、大数据与IT运维。
  • 识别AIOPS中的核心挑战,包括模型可解释性、不确定性评估以及自主修复能力。
  • 推动社区范围内的开放、可复现的研究实践,以加速AIOPS领域的创新。
  • 建立共享的基准测试框架,以实现AIOPS方法的跨比较,并跟踪研究进展。
  • 通过改进根本原因分析与自愈系统,推动实现完全自主的IT运维。

提出的方法

  • 对1,000多篇AIOPS相关文献进行了系统性映射研究,按宏观领域、类别及时间趋势进行分类。
  • 提出了AIOPS应用的分类体系,其中故障管理为最主要类别,进一步细分为故障检测、预测、根本原因分析与修复。
  • 采用因果建模技术,如霍克斯过程和图节点嵌入,以推断告警因果关系并识别根本原因告警。
  • 应用深度学习与服务依赖图,以高精度(0.92)检测云原生微服务中性能下降的根本原因。
  • 开发了一种去中心化的联邦学习方法,利用教师-学生蒸馏技术共享异常检测模型,而无需暴露原始日志数据或模型参数。
  • 提出了一种人工蜂群智能框架,用于公共云中的资源共享,以在保障用户体验质量(QoE)的同时优化资源利用率。

实验结果

研究问题

  • RQ1在时间和不同领域中,AIOPS的主导应用领域与研究趋势是什么?
  • RQ2因果推断与基于图的建模如何提升复杂IT系统中的根本原因分析能力?
  • RQ3在实现可信、可解释且自主的AI驱动IT运维过程中,面临哪些关键挑战?
  • RQ4如何在不暴露敏感数据的前提下,实现跨IT服务的隐私保护型知识共享?
  • RQ5公共基准在实现AIOPS研究中公平比较与加速进展方面发挥什么作用?

主要发现

  • 故障管理占所有AIOPS文献的62.1%,其中故障检测(33.7%)与根本原因分析(26.7%)为最活跃的子领域。
  • 近年来故障检测研究显著增长,2018–2019年间发表71篇,超过整个资源调配类别的总和(69篇)。
  • 基于深度学习的微服务性能诊断方法,结合服务依赖图与指标异常检测,实现了0.92的精确度,有效识别根本原因。
  • 联邦学习方法在不共享训练数据或模型参数的前提下,提升了基于日志的异常检测性能,有效保护了隐私。
  • 人工蜂群智能模型在高峰负载期间,有效平衡了云资源在各客户间的分配,在最小化资源利用率的同时维持了QoE。
  • 尽管关注度持续上升,目前仍缺乏用于AIOPS方法跨比较的公共基准,严重制约了该领域的进展与可复现性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。