Skip to main content
QUICK REVIEW

[Paper Review] Explainability Pitfalls: Beyond Dark Patterns in Explainable AI

Upol Ehsan, Mark Riedl|arXiv (Cornell University)|Sep 26, 2021
Explainable Artificial Intelligence (XAI)32 references17 citations
TL;DR

This paper introduces 'explainability pitfalls' (EPs)—unintended, negative downstream effects of AI explanations that arise even without malicious intent, distinct from dark patterns. Through a case study, it identifies how numerical explanations can mislead users into misplaced trust, and proposes proactive strategies at research, design, and organizational levels to detect and prevent such pitfalls using seamful design and EP literacy programs.

ABSTRACT

To make Explainable AI (XAI) systems trustworthy, understanding harmful effects is just as important as producing well-designed explanations. In this paper, we address an important yet unarticulated type of negative effect in XAI. We introduce explainability pitfalls(EPs), unanticipated negative downstream effects from AI explanations manifesting even when there is no intention to manipulate users. EPs are different from, yet related to, dark patterns, which are intentionally deceptive practices. We articulate the concept of EPs by demarcating it from dark patterns and highlighting the challenges arising from uncertainties around pitfalls. We situate and operationalize the concept using a case study that showcases how, despite best intentions, unsuspecting negative effects such as unwarranted trust in numerical explanations can emerge. We propose proactive and preventative strategies to address EPs at three interconnected levels: research, design, and organizational.

Motivation & Objective

  • To identify and conceptualize explainability pitfalls (EPs), unanticipated negative effects of AI explanations that occur without intentional deception.
  • To differentiate EPs from dark patterns by emphasizing intentionality and lack of malice in EPs.
  • To operationalize EPs through a controlled case study on user perception of numerical vs. natural language explanations.
  • To propose proactive, multi-level strategies—research, design, and organizational—to detect and prevent EPs in XAI systems.
  • To expand the XAI design space by raising awareness of previously unarticulated intellectual blind spots in explanation safety.

Proposed method

  • Conducts a controlled user study comparing perception of three AI explanation types: natural language with justification, natural language without justification, and uncontextualized numerical explanations.
  • Analyzes qualitative user responses to identify patterns of misinterpretation, such as misplaced trust in numerical outputs.
  • Uses speculative and reflective design methods to simulate potential pitfalls and envision 'what could go wrong' in XAI systems.
  • Proposes 'seamful explanations'—strategically revealing system mechanisms while concealing distractions—to promote reflective thinking and reduce reliance on misleading numerical cues.
  • Introduces 'pitfall literacy' programs for designers and users, using participatory design and simulation exercises to build awareness of EPs.
  • Advocates for organizational-level training and evaluation frameworks to embed EP awareness into XAI development and evaluation pipelines.

Experimental results

Research questions

  • RQ1How do users with and without AI backgrounds perceive uncontextualized numerical explanations in AI systems?
  • RQ2What types of unintended negative effects emerge from AI explanations even when designers have no intention to mislead?
  • RQ3In what ways do numerical explanations lead to misplaced trust or cognitive misalignment in users?
  • RQ4How can seamful design principles be applied to mitigate explainability pitfalls?
  • RQ5What organizational and educational strategies can improve resilience against unintended negative effects of explanations?

Key findings

  • Users without AI backgrounds exhibited significantly higher levels of misplaced trust in numerical explanations compared to natural language explanations, despite the numbers being uncontextualized.
  • Even without intentional deception, numerical explanations led users to overestimate the AI’s reliability and decision-making transparency.
  • The case study revealed that users often misinterpreted numerical outputs as inherently objective or authoritative, leading to over-reliance on them.
  • Seamful explanations—particularly interactive counterfactuals—were shown to promote reflective thinking and reduce cognitive shortcuts that lead to pitfalls.
  • Participants in the study demonstrated increased awareness of system limitations when exposed to contrastive 'what-if' scenarios, indicating that such methods can mitigate EPs.
  • The study underscores the need for EP literacy programs to equip both designers and users with tools to recognize and avoid unanticipated negative effects of explanations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.