Skip to main content
QUICK REVIEW

[论文解读] Assembling a Cyber Range to Evaluate Artificial Intelligence / Machine Learning (AI/ML) Security Tools

Jeffrey A. Nichols, Kevin D. Spakes|arXiv (Cornell University)|Jan 20, 2022
Network Security and Intrusion Detection被引用 4
一句话总结

本文介绍了在橡树岭国家实验室开发的可扩展、可编程网络攻防演练平台,通过可重复、受控的实验评估基于人工智能/机器学习(AI/ML)的网络安全工具。该平台能够自动化、高保真地模拟大规模网络环境——在两次国家级挑战中得到验证,涉及10万份文件样本和多阶段网络攻击——为AI/ML工具在真实、政府规模网络中的性能与运行成本提供标准化、可重复的评估。

ABSTRACT

In this case study, we describe the design and assembly of a cyber security testbed at Oak Ridge National Laboratory in Oak Ridge, TN, USA. The range is designed to provide agile reconfigurations to facilitate a wide variety of experiments for evaluations of cyber security tools -- particularly those involving AI/ML. In particular, the testbed provides realistic test environments while permitting control and programmatic observations/data collection during the experiments. We have designed in the ability to repeat the evaluations, so additional tools can be evaluated and compared at a later time. The system is one that can be scaled up or down for experiment sizes. At the time of the conference we will have completed two full-scale, national, government challenges on this range. These challenges are evaluating the performance and operating costs for AI/ML-based cyber security tools for application into large, government-sized networks. These evaluations will be described as examples providing motivation and context for various design decisions and adaptations we have made. The first challenge measured end-point security tools against 100K file samples (benignware and malware) chosen across a range of file types. The second is an evaluation of network intrusion detection systems efficacy in identifying multi-step adversarial campaigns -- involving reconnaissance, penetration and exploitations, lateral movement, etc. -- with varying levels of covertness in a high-volume business network. The scale of each of these challenges requires automation systems to repeat, or simultaneously mirror identical the experiments for each ML tool under test. Providing an array of easy-to-difficult malicious activity for sussing out the true abilities of the AI/ML tools has been a particularly interesting and challenging aspect of designing and executing these challenge events.

研究动机与目标

  • 设计并部署一个灵活、可重构的网络攻防演练平台,用于在真实网络环境中评估基于AI/ML的网络安全工具。
  • 通过程序化数据收集和可重复实验,实现对AI/ML工具的公平、标准化基准测试。
  • 支持在多样化威胁场景(包括多阶段网络攻击)下对AI/ML工具进行大规模、高容量的评估。
  • 通过自动化与模块化架构,降低运维开销并提高可重复性。
  • 为国家级网络安全挑战提供可扩展的测试平台,涵盖政府规模的网络和真实世界威胁模型。

提出的方法

  • 利用虚拟化和容器化环境构建网络攻防演练平台,以模拟具有可配置拓扑的企业级网络规模环境。
  • 支持对网络流量、系统状态和威胁注入的程序化控制,以实现可重复的实验。
  • 系统集成自动化数据采集管道,用于记录评估期间工具的行为、检测率和性能指标。
  • 使用合成但逼真的数据集(如涵盖良性软件和恶意软件的10万份文件样本)以实现一致的基准测试。
  • 架构支持在多个AI/ML工具之间并行执行相同实验,实现直接对比。
  • 测试平台可根据实验需求动态扩展或收缩,支持从小规模验证到大规模国家级挑战的各类场景。

实验结果

研究问题

  • RQ1如何设计网络攻防演练平台的体系结构,以支持在真实网络环境中对AI/ML安全工具进行可重复、大规模的评估?
  • RQ2哪些系统设计模式能够实现敏捷的重新配置和程序化控制,以适应多样化的AI/ML评估场景?
  • RQ3如何生成并扩展高保真、多阶段的网络攻击模拟,以评估AI/ML工具的检测能力?
  • RQ4需要哪些自动化机制,才能在相同条件下对多个AI/ML工具同时执行并行、一致的实验,以实现公平的基准测试?
  • RQ5如何在大规模、政府规模的网络模拟中,一致地衡量AI/ML工具的运行成本与性能?

主要发现

  • 该网络攻防演练平台成功支持了两次全规模国家级挑战,验证了其可扩展性与可重用性。
  • 系统实现了对多个AI/ML工具的可重复、程序化实验执行,并确保了数据收集的一致性。
  • 在10万份文件样本上对终端安全工具的评估,为恶意软件检测性能提供了标准化基准。
  • 网络入侵检测挑战模拟了复杂、多步骤的对抗性攻击活动,涵盖不同程度的隐蔽性,用以测试AI/ML工具的检测极限。
  • 自动化基础设施支持在相同条件下对多个AI/ML工具进行并行执行与对比,确保了结果的公平性与可重复性。
  • 测试平台的模块化设计使其能够快速重新配置并适应新的威胁模型与评估需求。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。