[Paper Review] Attacks, Defenses, And Tools: A Framework To Facilitate Robust AI/ML Systems
This paper proposes a comprehensive, systematically derived framework to catalog attacks, defenses, and tools in AI/ML systems, enabling proactive risk modeling and resilience planning. Based on a systematic literature review of 2,500 papers, it structures adversarial threats, mitigation techniques, and offensive/defensive tools with standardized attributes, forming a scalable online knowledge base for developers and researchers to assess and secure AI-enabled software systems.
Software systems are increasingly relying on Artificial Intelligence (AI) and Machine Learning (ML) components. The emerging popularity of AI techniques in various application domains attracts malicious actors and adversaries. Therefore, the developers of AI-enabled software systems need to take into account various novel cyber-attacks and vulnerabilities that these systems may be susceptible to. This paper presents a framework to characterize attacks and weaknesses associated with AI-enabled systems and provide mitigation techniques and defense strategies. This framework aims to support software designers in taking proactive measures in developing AI-enabled software, understanding the attack surface of such systems, and developing products that are resilient to various emerging attacks associated with ML. The developed framework covers a broad spectrum of attacks, mitigation techniques, and defensive and offensive tools. In this paper, we demonstrate the framework architecture and its major components, describe their attributes, and discuss the long-term goals of this research.
Motivation & Objective
- To address the lack of organized, systematic knowledge on AI/ML-specific cyber threats, vulnerabilities, and defenses.
- To create a structured, extensible framework that maps attacks, mitigation techniques, and tools to support proactive threat modeling.
- To characterize adversaries by goals, capabilities, assumptions, and expertise to improve threat intelligence in AI systems.
- To provide a scalable, online knowledge base that supports software designers in identifying attack surfaces and selecting appropriate defenses.
- To establish a foundation for long-term cyber-risk analysis of AI-enabled software systems through systematic categorization of adversarial techniques and tools.
Proposed method
- Conducted a systematic literature review across IEEE Xplore and ACM Digital Library (2000–2020), filtering 14,500 papers to 2,500 via title and abstract screening.
- Performed detailed analysis of selected papers and applied snowballing to identify additional relevant works not in the initial databases.
- Developed a meta-model with three core components: attacks, mitigation techniques, and tools, each with standardized attributes.
- Defined attributes for attacks (e.g., goal, specificity, assumption, capabilities), mitigation techniques (e.g., approach, type, advantage/disadvantage), and tools (e.g., input/output, offensive/defensive, availability).
- Established bidirectional mappings between attacks, mitigations, and tools to enable cross-referencing and informed decision-making.
- Implemented the framework as an online, filterable knowledge base at www.design.se.rit.edu/programs/ai-ml-framework for practical use by practitioners and researchers.
Experimental results
Research questions
- RQ1What are the key categories and attributes of adversarial attacks on AI/ML systems, and how can they be systematically classified?
- RQ2How do different mitigation techniques vary in effectiveness, cost, and deployment timing (proactive vs. reactive) across attack types?
- RQ3What offensive and defensive tools are currently available, and how do they relate to specific attacks and defense strategies?
- RQ4How can adversary models be formally characterized in terms of goals, capabilities, assumptions, and expertise to improve threat modeling?
- RQ5How can a unified, scalable framework integrate attacks, defenses, and tools to support cyber-risk analysis in AI-enabled software systems?
Key findings
- The framework successfully catalogs a broad spectrum of AI/ML attacks, including evasion, poisoning, membership inference, and model extraction, with standardized attributes such as goal, specificity, and assumption.
- Mitigation techniques are categorized as proactive (e.g., adversarial training) or reactive (e.g., adversarial sample detection), with documented advantages and disadvantages for each.
- A diverse set of offensive and defensive tools—ranging from open-source libraries like Foolbox and ART to commercial solutions—were identified and mapped to specific attacks and defenses.
- Adversary modeling reveals distinct threat profiles, including targeted vs. indiscriminate attacks and error-specific vs. error-generic goals, with varying levels of expertise and capabilities.
- The framework enables bi-directional mapping between attacks, mitigations, and tools, supporting informed selection of defense strategies based on threat context and system requirements.
- The initial version of the online knowledge base (www.design.se.rit.edu/programs/ai-ml-framework) is operational and supports filtering by attack type, mitigation, tool, and adversary profile, facilitating practical threat modeling.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.