[Paper Review] Sibyl: Understanding and Addressing the Usability Challenges of Machine Learning In High-Stakes Decision Making
This paper introduces Sibyl, a visual analytics tool designed to improve the usability of machine learning in high-stakes child welfare decision-making by making local factor contributions interpretable and interactive. Through iterative design with domain experts and two formal user studies, Sibyl effectively mitigates key usability challenges—such as trust, human-ML disagreement, and ethical concerns—by providing intuitive, case-specific explanations that enhance decision transparency and user confidence.
Machine learning (ML) is being applied to a diverse and ever-growing set of domains. In many cases, domain experts - who often have no expertise in ML or data science - are asked to use ML predictions to make high-stakes decisions. Multiple ML usability challenges can appear as result, such as lack of user trust in the model, inability to reconcile human-ML disagreement, and ethical concerns about oversimplification of complex problems to a single algorithm output. In this paper, we investigate the ML usability challenges that present in the domain of child welfare screening through a series of collaborations with child welfare screeners. Following the iterative design process between the ML scientists, visualization researchers, and domain experts (child screeners), we first identified four key ML challenges and honed in on one promising explainable ML technique to address them (local factor contributions). Then we implemented and evaluated our visual analytics tool, Sibyl, to increase the interpretability and interactivity of local factor contributions. The effectiveness of our tool is demonstrated by two formal user studies with 12 non-expert participants and 13 expert participants respectively. Valuable feedback was collected, from which we composed a list of design implications as a useful guideline for researchers who aim to develop an interpretable and interactive visualization tool for ML prediction models deployed for child welfare screeners and other similar domain experts.
Motivation & Objective
- To identify and address usability challenges faced by non-expert domain experts when using machine learning models for high-stakes decisions.
- To investigate how explainable AI (XAI) techniques, particularly local factor contributions, can improve interpretability and usability in qualitative decision-making domains.
- To co-design a visual analytics tool with child welfare screeners to support transparent, accountable, and ethically sound decision-making.
- To evaluate the effectiveness of the tool through formal user studies with both non-expert and expert participants.
- To derive actionable design implications for future XAI tools in similar high-stakes, non-technical domains.
Proposed method
- Conducted a comprehensive literature review of 55 papers on ML and explainability to identify recurring usability challenges in human-ML interaction.
- Collaborated iteratively with child welfare screeners, ML scientists, and visualization researchers to co-design Sibyl, focusing on local factor contributions as a core explanation technique.
- Implemented Sibyl as a visual analytics platform that displays case-specific feature contributions, similar cases, and counterfactual reasoning to support decision-making.
- Designed the interface with attention to cognitive biases, including sorting order of factor contributions to reduce availability bias.
- Conducted two formal user studies: one with 12 non-expert participants and another with 13 expert participants, using synthetic data to evaluate usability and perceived usefulness.
- Collected qualitative feedback and confidence ratings to assess the impact of different explanation types on user trust and decision clarity.
Experimental results
Research questions
- RQ1What are the primary usability challenges that non-expert domain experts face when using machine learning models in high-stakes, qualitative decision-making contexts?
- RQ2How can local factor contributions be visualized to effectively support human decision-making in child welfare screening?
- RQ3To what extent do interactive visual explanations improve user trust, reduce human-ML disagreement, and enhance perceived accountability?
- RQ4How do cognitive biases such as representativeness, causation vs. correlation, and availability bias influence user interpretation of ML explanations?
- RQ5What design principles emerge from user feedback that can guide the development of future XAI tools for non-technical, high-stakes domains?
Key findings
- Users exhibited low interest in understanding model mechanics, with only one screener expressing curiosity about model internals, indicating that transparency is less critical than actionable, case-specific insights.
- The local factor contribution explanation was perceived as the most useful tool for reconciling human-ML disagreements and building trust, outperforming global explanations and performance metrics.
- Sorting factor contributions in descending order (positive first) reduced the risk of availability bias compared to ascending order, which highlighted negative factors first and skewed perception.
- Users were more likely to misinterpret counterfactual explanations as causal relationships, highlighting the risk of reinforcing correlation-with-causation fallacies.
- The Similar Cases feature risked encouraging representativeness bias by prompting decisions based on case similarity rather than individual case features.
- Despite the synthetic data used, participants reported high perceived usefulness of Sibyl, particularly for understanding prediction drivers and justifying decisions in complex, ethically sensitive contexts.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.