[Paper Review] In AI We Trust? Factors That Influence Trustworthiness of AI-infused Decision-Making Processes
This study evaluates how seven factors (stakes, decider, trainer, interpretability, train/test description, social transparency, and model confidence) shape trust in AI-infused decisions, finding that interpretable models and transparent training data boost trust.
Many decision-making processes have begun to incorporate an AI element, including prison sentence recommendations, college admissions, hiring, and mortgage approval. In all of these cases, AI models are being trained to help human decision makers reach accurate and fair judgments, but little is known about what factors influence the extent to which people consider an AI-infused decision-making process to be trustworthy. We aim to understand how different factors about a decision-making process, and an AI model that supports that process, influences peoples' perceptions of the trustworthiness of that process. We report on our evaluation of how seven different factors -- decision stakes, decision authority, model trainer, model interpretability, social transparency, and model confidence -- influence ratings of trust in a scenario-based study.
Motivation & Objective
- Identify which factors about an AI-infused decision-making process influence perceived trustworthiness.
- Quantify the impact of each factor on multiple facets of trust (trustworthiness, reliability, technical competence, understandability, personal attachment).
- Examine interactions among factors (e.g., decider, trainer, confidence) and their effect on trust.
- Assess inter-individual variability in trust perceptions across crowd-sourced participants.
- Provide guidance for designing trusted AI systems by highlighting information needs of end users.
Proposed method
- Scenario-based evaluation with 128 unique combinations of factor levels (2x2x2x2x2x2x2).
- Crowdsourced rating study using Mechanical Turk with N=320 participants (two scenarios per participant, low and high stakes).
- Multi-faceted trust measurement using adapted scales for trustworthiness, reliability, technical competence, and personal attachment (4-point Likert scales).
- ANOVA with participant as a random effect to assess main effects and interactions, reporting partial eta-squared as effect size.
- Qualitative analysis of open-ended responses to understand reasons behind trust judgments.
Experimental results
Research questions
- RQ1How do the seven scenario factors influence different facets of trust in AI-infused decision making?
- RQ2Do interactions among factors (notably decider, trainer, and model confidence) modulate trust?
- RQ3What is the role of stakeholder perceptions and individual differences in trust judgments?
- RQ4Is interpretability and transparency in training data associated with higher trust across high- and low-stakes decisions?
Key findings
- Interpretable models consistently yield higher ratings for trustworthiness, reliability, and technical competence than black-box models.
- Providing information about how the model was trained and tested significantly increases trust-related ratings.
- Social transparency (information about decisions affecting others) has a smaller but positive effect on trust facets.
- Lower-stakes scenarios produce higher trust in AI-infused processes than higher-stakes ones.
- When AI is the decider, training by human data scientists (vs automated AI) increases trust-related judgments; when the decider is human, trainer effects are smaller.
- Model confidence visibility increases trust, especially when the decider is AI; absence of confidence information reduces trust.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.