Skip to main content
QUICK REVIEW

[Paper Review] Moral Foundations of Large Language Models

Marwa Abdulhai, Gregory Serapio‐García|arXiv (Cornell University)|Oct 23, 2023
Hate Speech and Cyberbullying Detection4 citations
TL;DR

This paper applies Moral Foundations Theory (MFT) to analyze moral biases in large language models (LLMs), revealing that LLMs exhibit consistent, ideology-linked moral foundations—such as a bias toward care/harm or loyalty—shaped by training data. It demonstrates that prompting can systematically alter these moral weights, significantly affecting downstream behavior, such as reducing charitable donations by 39% when harm is emphasized.

ABSTRACT

Moral foundations theory (MFT) is a psychological assessment tool that decomposes human moral reasoning into five factors, including care/harm, liberty/oppression, and sanctity/degradation (Graham et al., 2009). People vary in the weight they place on these dimensions when making moral decisions, in part due to their cultural upbringing and political ideology. As large language models (LLMs) are trained on datasets collected from the internet, they may reflect the biases that are present in such corpora. This paper uses MFT as a lens to analyze whether popular LLMs have acquired a bias towards a particular set of moral values. We analyze known LLMs and find they exhibit particular moral foundations, and show how these relate to human moral foundations and political affiliations. We also measure the consistency of these biases, or whether they vary strongly depending on the context of how the model is prompted. Finally, we show that we can adversarially select prompts that encourage the moral to exhibit a particular set of moral foundations, and that this can affect the model's behavior on downstream tasks. These findings help illustrate the potential risks and unintended consequences of LLMs assuming a particular moral stance.

Motivation & Objective

  • To investigate whether large language models (LLMs) inherit and reflect specific moral foundations from internet-scale training data.
  • To assess the consistency of these moral foundations across diverse conversational prompts and contexts.
  • To evaluate whether adversarial prompting can manipulate LLMs to adopt particular moral stances, such as those associated with liberal or conservative political ideologies.
  • To measure the downstream behavioral impact of such moral manipulations, particularly on tasks like charitable donation decisions.
  • To highlight ethical risks of LLMs amplifying or enforcing political or moral biases in real-world applications.

Proposed method

  • Adopted the Moral Foundations Questionnaire (MFQ), a 30-question inventory, to score LLMs on five moral dimensions: care/harm, fairness/cheating, loyalty/betrayal, authority/subversion, and sanctity/degradation.
  • Compared LLM responses to MFQ with human psychological studies to identify similarities and differences in moral foundation weighting.
  • Conducted consistency analysis by prompting the same LLM with diverse conversational contexts to test stability of moral foundation scores.
  • Designed adversarial prompts to induce LLMs to emphasize specific moral foundations, such as care/harm or loyalty.
  • Evaluated the behavioral impact of these moral prompts on a downstream dialog-based charitable donation benchmark.
  • Quantified the effect of moral foundation weighting on donation behavior, measuring differences in donation amounts across prompts.

Experimental results

Research questions

  • RQ1Do large language models exhibit consistent moral foundations derived from their training data, as measured by Moral Foundations Theory?
  • RQ2How stable are these moral foundations across different conversational prompts and contexts?
  • RQ3To what extent can adversarial prompting manipulate an LLM to adopt a specific moral foundation or political ideology?
  • RQ4Does shifting a model’s moral foundation through prompting lead to measurable changes in behavior on downstream tasks?
  • RQ5What are the ethical implications of LLMs exhibiting and being manipulatable into adopting particular moral stances?

Key findings

  • LLMs exhibit a consistent bias toward specific moral foundations, with GPT-3 showing a notable emphasis on care/harm and fairness, aligning with liberal political ideologies.
  • The moral foundation scores of LLMs remain relatively stable across diverse conversational prompts, indicating consistent internal moral framing.
  • Adversarial prompting can successfully shift LLMs to emphasize specific moral foundations, such as loyalty/betrayal or sanctity/degradation, mimicking conservative political stances.
  • When prompted to prioritize the harm foundation, LLMs donated 39% less in a charitable donation task compared to when prompted to emphasize loyalty.
  • The moral bias in LLMs extends beyond questionnaire responses and actively influences behavior on downstream tasks, indicating real-world behavioral consequences.
  • Fine-tuned LLMs, particularly those using reinforcement learning for safety, show reduced sensitivity in MFQ responses, leading to less human-like distributions and complicating bias measurement.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.