Skip to main content
QUICK REVIEW

[Paper Review] Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance

Kaitlyn Zhou, Jena D. Hwang|arXiv (Cornell University)|Jul 10, 2024
Complex Systems and Decision MakingDecision Sciences3 citations
TL;DR

This paper introduces Rel-A.I., an in situ, system-level evaluation framework that measures human reliance on language model (LM) outputs by analyzing epistemic markers (e.g., 'I think it's...') within dynamic interaction contexts. It reveals that reliance is significantly influenced by interactional factors—such as prior interactions, anthropomorphic cues, and subject domain—resulting in up to 20% variation in reliance frequency for identical expressions depending on context, challenging the assumption that verbalized confidence alone drives reliance.

ABSTRACT

The ability to communicate uncertainty, risk, and limitation is crucial for the safety of large language models. However, current evaluations of these abilities rely on simple calibration, asking whether the language generated by the model matches appropriate probabilities. Instead, evaluation of this aspect of LLM communication should focus on the behaviors of their human interlocutors: how much do they rely on what the LLM says? Here we introduce an interaction-centered evaluation framework called Rel-A.I. (pronounced "rely"}) that measures whether humans rely on LLM generations. We use this framework to study how reliance is affected by contextual features of the interaction (e.g, the knowledge domain that is being discussed), or the use of greetings communicating warmth or competence (e.g., "I'm happy to help!"). We find that contextual characteristics significantly affect human reliance behavior. For example, people rely 10% more on LMs when responding to questions involving calculations and rely 30% more on LMs that are perceived as more competent. Our results show that calibration and language quality alone are insufficient in evaluating the risks of human-LM interactions, and illustrate the need to consider features of the interactional context.

Motivation & Objective

  • To address the gap in understanding how interactional context influences human reliance on language models beyond verbalized confidence.
  • To develop a system-level evaluation method that captures in situ reliance behaviors during real human-LM interactions.
  • To investigate how factors like prior interactions, model persona (warmth), and subject domain affect reliance decisions.
  • To provide researchers and developers with a deployable methodology for assessing reliance before public deployment of language models.
  • To challenge the assumption that linguistic confidence alone determines reliance, emphasizing contextual and relational factors in human-AI interaction.

Proposed method

  • Rel-A.I. employs a self-incentivized task design where users engage in naturalistic, multi-turn interactions with LMs to measure reliance in real-time.
  • It identifies and tracks epistemic markers (e.g., 'I believe', 'Maybe', 'Undoubtedly') in LM-generated responses as proxies for confidence and uncertainty.
  • The method incorporates meta-level perception questions assessing users’ perceived warmth and competence of the AI agent to capture anthropomorphic influences.
  • It evaluates reliance across three interaction settings: long-term interactions, anthropomorphic generations, and variable subject matter.
  • Reliance is measured as the frequency with which users adopt or echo LM-generated epistemic markers in their own responses.
  • The framework enables in situ, system-level measurement of reliance without relying on post-hoc self-reports or synthetic confidence annotations.

Experimental results

Research questions

  • RQ1How do prior interactions with a language model influence human reliance on its outputs, particularly when confidence levels vary?
  • RQ2To what extent does the perceived warmth or anthropomorphic persona of a language model affect human reliance on its responses?
  • RQ3How does the subject domain of a conversation (e.g., philosophy, law, religion) influence reliance on LM-generated epistemic markers?
  • RQ4Does the same expression of moderate confidence ('I'm pretty sure it's...') lead to different reliance rates depending on the interaction context?
  • RQ5How do contextual features such as model consistency and perceived personality interact with linguistic confidence to shape reliance behavior?

Key findings

  • Reliance on language model outputs varies by up to 20% for the same epistemic expression depending on the interactional context, such as prior model behavior.
  • Statements of moderate confidence ('I'm pretty sure it's...') are relied on significantly less when generated by a model that is typically highly confident, compared to one that is usually uncertain.
  • Anthropomorphic cues, such as 'I'm happy to help!', significantly increase reliance on the same factual claim, even when the confidence level is identical.
  • Subject domain significantly affects reliance, with higher reliance observed in domains like philosophy and religion compared to computationally heavy domains like math.
  • The perception of warmth and competence in the AI agent independently influences reliance, even when linguistic confidence is held constant.
  • Verbalized confidence alone is insufficient to predict reliance; interactional context—including history, persona, and topic—must be considered in human-LM reliance models.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.