Skip to main content
QUICK REVIEW

[论文解读] Who's Thinking? A Push for Human-Centered Evaluation of LLMs using the XAI Playbook

Teresa Datta, John P. Dickerson|arXiv (Cornell University)|Mar 10, 2023
Topic Modeling被引用 4
一句话总结

本文通过借鉴可解释人工智能(XAI)的实践方法,倡导以人类为中心的大语言模型(LLM)评估,强调心智模型、应用场景实用性以及认知参与度。文章认为,LLM 必须通过人类认知、偏见和现实世界任务来评估,而不仅仅是技术指标,以突出过度依赖和幻觉带来的风险。

ABSTRACT

Deployed artificial intelligence (AI) often impacts humans, and there is no one-size-fits-all metric to evaluate these tools. Human-centered evaluation of AI-based systems combines quantitative and qualitative analysis and human input. It has been explored to some depth in the explainable AI (XAI) and human-computer interaction (HCI) communities. Gaps remain, but the basic understanding that humans interact with AI and accompanying explanations, and that humans' needs -- complete with their cognitive biases and quirks -- should be held front and center, is accepted by the community. In this paper, we draw parallels between the relatively mature field of XAI and the rapidly evolving research boom around large language models (LLMs). Accepted evaluative metrics for LLMs are not human-centered. We argue that many of the same paths tread by the XAI community over the past decade will be retread when discussing LLMs. Specifically, we argue that humans' tendencies -- again, complete with their cognitive biases and quirks -- should rest front and center when evaluating deployed LLMs. We outline three developed focus areas of human-centered evaluation of XAI: mental models, use case utility, and cognitive engagement, and we highlight the importance of exploring each of these concepts for LLMs. Our goal is to jumpstart human-centered LLM evaluation.

研究动机与目标

  • 解决尽管 LLM 广泛用于公众场景,但其评估仍缺乏以人类为中心的问题。
  • 识别当前 LLM 评估框架中忽略人类认知、偏见和现实世界决策制定的缺陷。
  • 将可解释人工智能(XAI)的洞见迁移至 LLM,主张应以类似的人类中心评估原则指导 LLM 的开发。
  • 突出用户对 LLM 输出缺乏审视的认知参与可能带来的风险,包括过度信任和幻觉现象。
  • 呼吁开展研究,探讨用户如何形成对 LLM 的心智模型,其在实际任务中的实用性如何,以及用户与输出的参与度如何。

提出的方法

  • 将 XAI 评估框架——特别是心智模型、应用场景实用性与认知参与度——转化为 LLM 评估的结构化方法。
  • 在 XAI 的人类中心评估与当前 LLM 状态之间建立类比,强调在信任、可解释性与可用性方面面临的共同挑战。
  • 提出基于应用场景的用户研究,以评估 LLM 在现实世界任务(如邮件撰写或决策支持)中的实用性。
  • 引入认知强制策略(例如预检任务),以提升用户参与度,减少对 LLM 的过度依赖所导致的错误。
  • 采用定性与定量分析,评估用户如何理解、信任并互动于 LLM 生成的内容。
  • 强调不仅需衡量模型性能,还应关注人机交互动态,包括确认偏误和与用户意图的错位。

实验结果

研究问题

  • RQ1普通用户如何形成对 LLM 生成响应机制的心智模型?
  • RQ2LLM 在现实应用场景(如邮件撰写或决策支持)中提供了多大程度的实际实用性?
  • RQ3用户对 LLM 输出的认知参与度在多大程度上影响其行为、信任度与错误率?
  • RQ4认知偏见(如确认偏误)在多大程度上影响用户对 LLM 生成内容的解读?
  • RQ5认知强制策略如何提升用户对 LLM 输出的审慎态度,减少过度依赖?

主要发现

  • 当前的 LLM 评估指标并非以人类为中心,未能考虑认知偏见、心智模型与现实世界可用性。
  • 正如在 XAI 中一样,LLM 的解释没有真实标准,因此仅靠技术上的忠实度与鲁棒性不足以完成评估。
  • 实践中,用户对 LLM 输出的认知参与度较低,尤其是在时间压力下或面对高置信度的幻觉时,错误风险增加。
  • 研究表明,用户即使在 LLM 输出错误时仍倾向于信任其结果,尤其当回应表达自信时,这种现象被类比为“以服务形式呈现的说教式表达”。
  • 认知强制策略(如要求用户在采纳前验证输出)已被证明能显著提升表现并减少错误。
  • LLM 部署规模,尤其是在公众使用的工具(如 ChatGPT)中,放大了未经审查的认知卸载风险,迫切需要以人类为中心的评估。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。