[Paper Review] Hows and Whys of Artificial Intelligence for Public Sector Decisions: Explanation and Evaluation
This paper proposes a framework for evaluating and explaining AI/ML systems in public sector decision-making by modeling human-AI decision loops and mapping evaluation and explanation challenges across four strategies: mission-, data-, work-, and evidence-oriented. It demonstrates that aligning AI evaluation with organizational outcomes and human-machine collaboration improves trust, accountability, and impact, offering actionable pathways for public sector adoption.
Evaluation has always been a key challenge in the development of artificial intelligence (AI) based software, due to the technical complexity of the software artifact and, often, its embedding in complex sociotechnical processes. Recent advances in machine learning (ML) enabled by deep neural networks has exacerbated the challenge of evaluating such software due to the opaque nature of these ML-based artifacts. A key related issue is the (in)ability of such systems to generate useful explanations of their outputs, and we argue that the explanation and evaluation problems are closely linked. The paper models the elements of a ML-based AI system in the context of public sector decision (PSD) applications involving both artificial and human intelligence, and maps these elements against issues in both evaluation and explanation, showing how the two are related. We consider a number of common PSD application patterns in the light of our model, and identify a set of key issues connected to explanation and evaluation in each case. Finally, we propose multiple strategies to promote wider adoption of AI/ML technologies in PSD, where each is distinguished by a focus on different elements of our model, allowing PSD policy makers to adopt an approach that best fits their context and concerns.
Motivation & Objective
- To address the persistent challenge of evaluating AI systems in public sector decision-making (PSD), where technical complexity and sociotechnical embedding complicate verification and validation.
- To clarify the interdependence between explanation and evaluation in AI/ML systems, especially in light of the opacity of deep learning models.
- To develop a structured model of human-AI decision loops that identifies key points of failure and opportunity for explanation and evaluation.
- To propose four distinct strategies—mission-, data-, work-, and evidence-oriented—for public sector organizations to adopt AI based on their specific priorities and constraints.
- To promote sustainable, trustworthy AI adoption by grounding evaluation in real-world impact, not just model accuracy, and by emphasizing feedback and learning loops.
Proposed method
- Develops a conceptual model of a human-AI decision loop adapted from Boyd’s OODA loop, distinguishing data elements (input, training, feedback) and operational elements (observe, orient, decide, act).
- Maps evaluation (verification and validation) and explanation (How? and Why?) to specific components of the decision loop, identifying where scrutiny is most critical.
- Proposes four distinct strategies for AI deployment in PSD: mission-oriented (focus on outcomes), data-oriented (focus on data quality and management), work-oriented (focus on human-machine collaboration), and evaluation-oriented (holistic, evidence-based approach).
- Uses the framework to analyze common PSD application patterns, identifying key issues in explanation and evaluation for each.
- Integrates concepts of single-, double-, and triple-loop learning to emphasize iterative improvement through feedback and evidence-based validation.
- Draws on historical lessons from AI winters to argue that long-term success depends on empirical evidence of impact, not just technical performance.
Experimental results
Research questions
- RQ1How can evaluation and explanation in AI/ML systems for public sector decisions be systematically linked to improve trust and accountability?
- RQ2What are the key differences in explanation needs between verification (‘building the system right’) and validation (‘building the right system’) in public sector AI applications?
- RQ3How do the four proposed strategies—mission-, data-, work-, and evidence-oriented—differentiate in their focus and impact on AI evaluation and explanation?
- RQ4In what ways do sociotechnical processes in public sector organizations complicate the evaluation of AI system impacts, and how can this be addressed?
- RQ5How can feedback loops and evidence-based evaluation help overcome the risk of AI hype and ensure sustainable, high-impact AI deployment in public services?
Key findings
- Explanation and evaluation are deeply intertwined: effective explanation supports both verification (e.g., debugging) and validation (e.g., trust), especially in complex sociotechnical systems.
- The mission-oriented strategy prioritizes decision, action, and feedback data, focusing explanation on 'why' outcomes succeeded or failed, thereby linking AI performance to real-world impact.
- The data-oriented strategy enhances evaluation by improving data quality and management, shifting the burden toward verification and increasing the 'known knowns' in the system.
- The work-oriented strategy improves systemic robustness and job satisfaction by optimizing the division of labor between humans and AI, with explanation playing a key role in trust and accountability.
- The evaluation-oriented strategy provides a holistic, evidence-based approach that captures 'what works' across the entire decision loop, supporting long-term learning and institutional memory.
- Feedback loops—especially triple-loop learning—are essential for adapting AI systems to novel contexts and mitigating risks from 'unknown knowns' and 'unknown unknowns', ensuring sustained impact beyond initial performance metrics.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.