Skip to main content
QUICK REVIEW

[Paper Review] Who's Thinking? A Push for Human-Centered Evaluation of LLMs using the XAI Playbook

Teresa Datta, John P. Dickerson|arXiv (Cornell University)|Mar 10, 2023
Topic Modeling4 citations
TL;DR

This paper advocates for human-centered evaluation of large language models (LLMs) by adapting the Explainable AI (XAI) playbook, emphasizing mental models, use case utility, and cognitive engagement. It argues that LLMs must be evaluated through human cognition, biases, and real-world tasks—not just technical metrics—highlighting risks of over-reliance and hallucination.

ABSTRACT

Deployed artificial intelligence (AI) often impacts humans, and there is no one-size-fits-all metric to evaluate these tools. Human-centered evaluation of AI-based systems combines quantitative and qualitative analysis and human input. It has been explored to some depth in the explainable AI (XAI) and human-computer interaction (HCI) communities. Gaps remain, but the basic understanding that humans interact with AI and accompanying explanations, and that humans' needs -- complete with their cognitive biases and quirks -- should be held front and center, is accepted by the community. In this paper, we draw parallels between the relatively mature field of XAI and the rapidly evolving research boom around large language models (LLMs). Accepted evaluative metrics for LLMs are not human-centered. We argue that many of the same paths tread by the XAI community over the past decade will be retread when discussing LLMs. Specifically, we argue that humans' tendencies -- again, complete with their cognitive biases and quirks -- should rest front and center when evaluating deployed LLMs. We outline three developed focus areas of human-centered evaluation of XAI: mental models, use case utility, and cognitive engagement, and we highlight the importance of exploring each of these concepts for LLMs. Our goal is to jumpstart human-centered LLM evaluation.

Motivation & Objective

  • Address the lack of human-centered evaluation in LLMs despite their widespread public use.
  • Identify gaps in current LLM evaluation frameworks that ignore human cognition, biases, and real-world decision-making.
  • Bridge insights from Explainable AI (XAI) to LLMs, arguing that similar human-centered evaluation principles should guide LLM development.
  • Highlight the risks of unexamined cognitive engagement with LLM outputs, including over-trust and hallucination.
  • Call for research that investigates how users form mental models of LLMs, how useful they are in real tasks, and how engaged users are with outputs.

Proposed method

  • Adapt the XAI evaluation framework—specifically mental models, use case utility, and cognitive engagement—into a structured approach for LLM evaluation.
  • Draw parallels between XAI’s human-centered evaluation and the current state of LLMs, emphasizing shared challenges in trust, interpretability, and usability.
  • Propose use case-based user studies to evaluate LLM utility in real-world tasks, such as email drafting or decision support.
  • Introduce cognitive forcing strategies (e.g., pre-check tasks) to increase user engagement and reduce errors from over-reliance on LLMs.
  • Use qualitative and quantitative analysis to assess how users interpret, trust, and interact with LLM-generated content.
  • Highlight the importance of measuring not just model performance, but also human-AI interaction dynamics, including confirmation bias and misalignment with user intent.

Experimental results

Research questions

  • RQ1How do everyday users form mental models of how LLMs generate responses?
  • RQ2To what extent do LLMs provide practical utility in real-world use cases such as email writing or decision support?
  • RQ3How does cognitive engagement with LLM outputs affect user behavior, trust, and error rates?
  • RQ4In what ways do cognitive biases—like confirmation bias—impact user interpretation of LLM-generated content?
  • RQ5How can cognitive forcing strategies improve user scrutiny of LLM outputs and reduce over-reliance?

Key findings

  • Current LLM evaluation metrics are not human-centered and fail to account for cognitive biases, mental models, and real-world usability.
  • Just as in XAI, there is no ground truth for explanations in LLMs, making technical faithfulness and robustness insufficient alone for evaluation.
  • Cognitive engagement with LLM outputs is low in practice, especially under time pressure or with high-confidence hallucinations, increasing error risk.
  • Studies show that users often trust LLMs even when outputs are incorrect, particularly when responses are confidently worded, a phenomenon likened to 'mansplaining as a service'.
  • Cognitive forcing strategies—such as requiring users to verify outputs before acceptance—have been shown to significantly improve performance and reduce errors.
  • The scale of LLM deployment, especially in public-facing tools like ChatGPT, amplifies risks of unchecked cognitive offloading, demanding urgent human-centered evaluation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.