[Paper Review] A.I. Robustness: a Human-Centered Perspective on Technological Challenges and Opportunities
This paper presents a human-centered framework for understanding AI robustness by unifying fragmented terminology across AI domains. It introduces three taxonomies—robustness by machine learning pipeline stages, model/task-specific robustness, and assessment methodologies—highlighting human involvement in evaluating and improving robustness, while identifying key research gaps and future directions for trustworthy AI systems.
Despite the impressive performance of Artificial Intelligence (AI) systems, their robustness remains elusive and constitutes a key issue that impedes large-scale adoption. Robustness has been studied in many domains of AI, yet with different interpretations across domains and contexts. In this work, we systematically survey the recent progress to provide a reconciled terminology of concepts around AI robustness. We introduce three taxonomies to organize and describe the literature both from a fundamental and applied point of view: 1) robustness by methods and approaches in different phases of the machine learning pipeline; 2) robustness for specific model architectures, tasks, and systems; and in addition, 3) robustness assessment methodologies and insights, particularly the trade-offs with other trustworthiness properties. Finally, we identify and discuss research gaps and opportunities and give an outlook on the field. We highlight the central role of humans in evaluating and enhancing AI robustness, considering the necessary knowledge humans can provide, and discuss the need for better understanding practices and developing supportive tools in the future.
Motivation & Objective
- To reconcile the inconsistent terminology and interpretations of AI robustness across different domains and applications.
- To identify and systematize key challenges in AI robustness that hinder large-scale deployment of AI systems.
- To emphasize the critical role of human expertise in evaluating and enhancing AI robustness beyond automated metrics.
- To map existing robustness methodologies across the machine learning lifecycle and specific model architectures.
- To highlight trade-offs between robustness and other trustworthiness properties such as fairness, interpretability, and efficiency.
Proposed method
- Systematic survey of recent literature on AI robustness across diverse domains, model architectures, and application contexts.
- Development of three interrelated taxonomies: (1) robustness by machine learning pipeline phase (data, training, inference), (2) robustness by model architecture and task, and (3) robustness assessment methodologies.
- Analysis of trade-offs between robustness and other AI trustworthiness properties, including fairness, interpretability, and efficiency.
- Incorporation of human-in-the-loop perspectives through case studies and expert insights on evaluating robustness.
- Identification of gaps in current robustness evaluation practices and the need for human-aware tools and frameworks.
- Synthesis of findings into a unified conceptual framework for guiding future research and development in robust AI.
Experimental results
Research questions
- RQ1How can inconsistent terminology and interpretations of AI robustness across domains be reconciled?
- RQ2What are the key robustness challenges specific to different stages of the machine learning pipeline?
- RQ3How does robustness vary across different model architectures and AI tasks?
- RQ4What are the trade-offs between robustness and other trustworthiness properties in AI systems?
- RQ5In what ways can human expertise improve the evaluation and enhancement of AI robustness?
Key findings
- The paper identifies significant terminological fragmentation in AI robustness research, with varying definitions across domains and applications.
- Robustness is most effectively addressed through a lifecycle perspective, with distinct challenges emerging in data preparation, model training, and inference stages.
- Model-specific robustness issues are prevalent in deep neural networks, especially in vision and NLP tasks, due to sensitivity to adversarial examples and distributional shifts.
- Assessment methodologies often overlook human perception and cognitive biases, leading to overestimation of robustness in real-world settings.
- A strong trade-off exists between robustness and model efficiency or interpretability, particularly in resource-constrained environments.
- Human involvement is essential not only for evaluating robustness but also for identifying edge cases and contextual failures that automated metrics miss.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.