Skip to main content
QUICK REVIEW

[Paper Review] MoralBench: Moral Evaluation of LLMs

Jianchao Ji, Yutong Chen|arXiv (Cornell University)|Jun 6, 2024
Legal Education and Practice Innovations7 citations
TL;DR

MoralBench introduces benchmark datasets and binary/comparative evaluation methods to measure LLMs’ moral identity aligned with Moral Foundations Theory, revealing model variability across domains.

ABSTRACT

In the rapidly evolving field of artificial intelligence, large language models (LLMs) have emerged as powerful tools for a myriad of applications, from natural language processing to decision-making support systems. However, as these models become increasingly integrated into societal frameworks, the imperative to ensure they operate within ethical and moral boundaries has never been more critical. This paper introduces a novel benchmark designed to measure and compare the moral reasoning capabilities of LLMs. We present the first comprehensive dataset specifically curated to probe the moral dimensions of LLM outputs, addressing a wide range of ethical dilemmas and scenarios reflective of real-world complexities. The main contribution of this work lies in the development of benchmark datasets and metrics for assessing the moral identity of LLMs, which accounts for nuance, contextual sensitivity, and alignment with human ethical standards. Our methodology involves a multi-faceted approach, combining quantitative analysis with qualitative insights from ethics scholars to ensure a thorough evaluation of model performance. By applying our benchmark across several leading LLMs, we uncover significant variations in moral reasoning capabilities of different models. These findings highlight the importance of considering moral reasoning in the development and evaluation of LLMs, as well as the need for ongoing research to address the biases and limitations uncovered in our study. We publicly release the benchmark at https://drive.google.com/drive/u/0/folders/1k93YZJserYc2CkqP8d4B3M3sgd3kA8W7 and also open-source the code of the project at https://github.com/agiresearch/MoralBench.

Motivation & Objective

  • Motivate the need to evaluate LLMs for moral alignment given their societal impact.
  • Develop two benchmark datasets grounded in Moral Foundations Theory to probe moral reasoning in LLM outputs.
  • Propose a binary and a comparative evaluation framework to quantify moral identity in LLMs.
  • Apply the benchmark to multiple LLMs to characterize strengths/limitations and guide ethical AI development.

Proposed method

  • Adopt Moral Foundations Theory (six foundations) to build two datasets: MFQ-30-LLM and MFV-LLM.
  • Use MFQ-30-LLM to obtain a two-part 0-5 scale assessment per foundation and aggregate into a total score.
  • Use MFV-LLM with 132 vignettes to generate foundation-specific wrongness ratings for evaluation.
  • Implement binary moral assessment where models respond Agree/Disagree and score via a human-reference mapping (binary + comparative scoring).
  • Implement comparative moral assessment where models choose the more moral statement between pairs, scored against higher human-average scores.

Experimental results

Research questions

  • RQ1Can LLMs consistently reflect human moral foundations across dimensions (Care/Harm, Fairness, Loyalty, Authority, Sanctity, Liberty)?
  • RQ2Do binary and comparative evaluation modes yield consistent or divergent indications of moral identity in LLMs?
  • RQ3Which models best align with human moral judgments across MFQ-LLM and MFV-LLM benchmarks?
  • RQ4What does discrepancy between binary and comparative results imply about models’ deep moral understanding?

Key findings

  • LLaMA-2 and GPT-4 achieve the highest performance in the binary MFQ-30-LLM and MFV-LLM benchmarks, respectively.
  • GPT-3.5 often scores highest overall in both MFQ-30-LLM and MFV-LLM comparative assessments, indicating strong ability to distinguish moral statements.
  • Gemma-1.1 consistently underperforms relative to other models across datasets and tasks.
  • Some models show high binary-moral-identity scores but struggle in comparative tasks, suggesting surface-level pattern recognition rather than deep moral understanding.
  • The benchmark reveals model-specific strengths across moral foundations and highlights variability in moral reasoning capabilities across architectures.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.