Skip to main content
QUICK REVIEW

[Paper Review] Assembling a Cyber Range to Evaluate Artificial Intelligence / Machine Learning (AI/ML) Security Tools

Jeffrey A. Nichols, Kevin D. Spakes|arXiv (Cornell University)|Jan 20, 2022
Network Security and Intrusion Detection4 citations
TL;DR

This paper presents a scalable, programmable cyber range developed at Oak Ridge National Laboratory to evaluate AI/ML-based cybersecurity tools through repeatable, controlled experiments. It enables automated, high-fidelity simulation of large-scale network environments—demonstrated in two national challenges involving 100K file samples and multi-stage cyber campaigns—providing standardized, repeatable assessments of AI/ML tool performance and operational costs in realistic, government-sized networks.

ABSTRACT

In this case study, we describe the design and assembly of a cyber security testbed at Oak Ridge National Laboratory in Oak Ridge, TN, USA. The range is designed to provide agile reconfigurations to facilitate a wide variety of experiments for evaluations of cyber security tools -- particularly those involving AI/ML. In particular, the testbed provides realistic test environments while permitting control and programmatic observations/data collection during the experiments. We have designed in the ability to repeat the evaluations, so additional tools can be evaluated and compared at a later time. The system is one that can be scaled up or down for experiment sizes. At the time of the conference we will have completed two full-scale, national, government challenges on this range. These challenges are evaluating the performance and operating costs for AI/ML-based cyber security tools for application into large, government-sized networks. These evaluations will be described as examples providing motivation and context for various design decisions and adaptations we have made. The first challenge measured end-point security tools against 100K file samples (benignware and malware) chosen across a range of file types. The second is an evaluation of network intrusion detection systems efficacy in identifying multi-step adversarial campaigns -- involving reconnaissance, penetration and exploitations, lateral movement, etc. -- with varying levels of covertness in a high-volume business network. The scale of each of these challenges requires automation systems to repeat, or simultaneously mirror identical the experiments for each ML tool under test. Providing an array of easy-to-difficult malicious activity for sussing out the true abilities of the AI/ML tools has been a particularly interesting and challenging aspect of designing and executing these challenge events.

Motivation & Objective

  • To design and deploy a flexible, reconfigurable cyber range for evaluating AI/ML-based cybersecurity tools in realistic network environments.
  • To enable programmatic data collection and repeatable experimentation for fair, standardized benchmarking of AI/ML tools.
  • To support large-scale, high-volume evaluations of AI/ML tools across diverse threat scenarios, including multi-stage cyber campaigns.
  • To reduce operational overhead and increase reproducibility through automation and modular architecture.
  • To provide a scalable testbed for national-level cybersecurity challenges involving government-sized networks and real-world threat profiles.

Proposed method

  • The cyber range is built using virtualized and containerized environments to simulate enterprise-scale networks with configurable topologies.
  • It supports programmatic control of network traffic, system states, and threat injection to enable repeatable experiments.
  • The system integrates automated data collection pipelines to log tool behavior, detection rates, and performance metrics during evaluations.
  • It uses synthetic yet realistic datasets—such as 100K file samples spanning benignware and malware—for consistent benchmarking.
  • The architecture supports parallel execution of identical experiments across multiple AI/ML tools to enable direct comparison.
  • The testbed is designed to scale up or down based on experiment needs, supporting both small-scale validation and large-scale national challenges.

Experimental results

Research questions

  • RQ1How can a cyber range be architected to support repeatable, large-scale evaluations of AI/ML-based security tools in realistic network environments?
  • RQ2What system design patterns enable agile reconfiguration and programmatic control for diverse AI/ML evaluation scenarios?
  • RQ3How can high-fidelity, multi-stage cyber attack simulations be generated and scaled for assessing AI/ML tool detection capabilities?
  • RQ4What automation mechanisms are required to support simultaneous, identical experiments across multiple AI/ML tools for fair benchmarking?
  • RQ5How can the operational cost and performance of AI/ML tools be measured consistently across large-scale, government-sized network simulations?

Key findings

  • The cyber range successfully supported two full-scale national challenges involving AI/ML tools, demonstrating its scalability and reusability.
  • The system enabled repeatable, programmatic execution of experiments with consistent data collection across multiple AI/ML tools.
  • The evaluation of endpoint security tools on 100K file samples provided a standardized benchmark for malware detection performance.
  • The network intrusion detection challenge simulated complex, multi-step adversarial campaigns with varying levels of covertness, testing the detection limits of AI/ML tools.
  • The automation infrastructure allowed for simultaneous execution and comparison of multiple AI/ML tools under identical conditions, ensuring fair and reproducible results.
  • The testbed’s modular design allowed for rapid reconfiguration and adaptation to new threat models and evaluation requirements.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.