Skip to main content
QUICK REVIEW

[Paper Review] Investigating Agency of LLMs in Human-AI Collaboration Tasks

Ashish Sharma, Sudha Rao|arXiv (Cornell University)|May 22, 2023
Speech and dialogue systems4 citations
TL;DR

This paper proposes a framework to measure and control agency in large language models (LLMs) during human-AI collaboration, grounded in social-cognitive theory. It identifies four agency features—Intentionality, Motivation, Self-Efficacy, and Self-Regulation—and introduces a new dataset of 83 human-human interior design dialogues (908 snippets) annotated for these features. Results show that models expressing high levels of these features are perceived as significantly more agentive in both automatic and human evaluations.

ABSTRACT

Agency, the capacity to proactively shape events, is central to how humans interact and collaborate. While LLMs are being developed to simulate human behavior and serve as human-like agents, little attention has been given to the Agency that these models should possess in order to proactively manage the direction of interaction and collaboration. In this paper, we investigate Agency as a desirable function of LLMs, and how it can be measured and managed. We build on social-cognitive theory to develop a framework of features through which Agency is expressed in dialogue - indicating what you intend to do (Intentionality), motivating your intentions (Motivation), having self-belief in intentions (Self-Efficacy), and being able to self-adjust (Self-Regulation). We collect a new dataset of 83 human-human collaborative interior design conversations containing 908 conversational snippets annotated for Agency features. Using this dataset, we develop methods for measuring Agency of LLMs. Automatic and human evaluations show that models that manifest features associated with high Intentionality, Motivation, Self-Efficacy, and Self-Regulation are more likely to be perceived as strongly agentive.

Motivation & Objective

  • To investigate how agency can be meaningfully expressed and measured in LLMs during human-AI collaboration.
  • To develop a theory-driven framework of agency features based on social-cognitive theory.
  • To collect and annotate a new dataset of human-human collaborative interior design dialogues for agency analysis.
  • To evaluate how well LLMs can express agency features and how these expressions affect perceived agentivity.
  • To enable control and measurement of agency levels in dialogue systems for improved human-AI collaboration.

Proposed method

  • Adopting Bandura’s social-cognitive theory, the authors define four core features of agency: Intentionality (expressing preferences), Motivation (justifying intentions), Self-Efficacy (asserting belief in judgments), and Self-Regulation (adapting based on feedback).
  • The authors collected 83 human-human collaborative interior design conversations, totaling 908 conversational snippets, and annotated each snippet for the four agency features.
  • They developed automatic measurement methods using fine-tuned LLMs to predict agency feature expressions from dialogue turns.
  • Two new evaluation tasks were introduced: (1) Measuring Agency in Dialogue, and (2) Generating Dialogue with Agency, to assess model performance.
  • Human evaluations were conducted to validate the automatic measurements and assess perceived agentivity across different agency feature levels.
  • The framework was evaluated using both automatic metrics and human judgments, showing strong correlation between expressed agency features and perceived agentivity.

Experimental results

Research questions

  • RQ1How can agency in LLMs be systematically defined and measured in dialogue-based human-AI collaboration?
  • RQ2To what extent do expressions of Intentionality, Motivation, Self-Efficacy, and Self-Regulation influence perceived agentivity in LLMs?
  • RQ3Can automatic methods reliably detect and quantify agency features in conversational turns?
  • RQ4How do different levels of agency expression affect human perception of LLMs as proactive collaborators?
  • RQ5Can dialogue systems be designed to dynamically modulate their agency based on context and user preferences?

Key findings

  • Models that express high levels of Intentionality, Motivation, Self-Efficacy, and Self-Regulation are consistently perceived as more agentive in both automatic and human evaluations.
  • Intentionality was found to be the strongest predictor of perceived agency in collaborative dialogue, significantly influencing overall agentive perception.
  • Human evaluations confirmed that agency features such as self-belief (Self-Efficacy) and adaptive behavior (Self-Regulation) contribute meaningfully to perceived proactivity in LLMs.
  • The proposed automatic measurement methods showed strong alignment with human judgments, indicating feasibility for scalable agency assessment.
  • The dataset and framework are adaptable beyond interior design, as core agency features like 'I would prefer' are domain-independent and transferable.
  • The study demonstrates that agency in LLMs is not just a passive trait but can be actively controlled and modulated for better human-AI collaboration.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.