Skip to main content
QUICK REVIEW

[Paper Review] LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models

Gwenyth Isobel Meadows, Nuno Lau|arXiv (Cornell University)|Jul 27, 2024
Hate Speech and Cyberbullying Detection4 citations
TL;DR

LocalValueBench introduces a collaboratively built, extensible benchmark to evaluate localized value alignment and ethical safety in large language models (LLMs), focusing on Australian values. Using a three-layered interrogation framework—baseline, debate, and forced justification—alongside human-validated scoring, the study reveals significant deviations in commercial LLMs (e.g., ChatGPT, Gemini, Claude) from local ethical standards, especially under adversarial prompting.

ABSTRACT

The proliferation of large language models (LLMs) requires robust evaluation of their alignment with local values and ethical standards, especially as existing benchmarks often reflect the cultural, legal, and ideological values of their creators. extsc{LocalValueBench}, introduced in this paper, is an extensible benchmark designed to assess LLMs' adherence to Australian values, and provides a framework for regulators worldwide to develop their own LLM benchmarks for local value alignment. Employing a novel typology for ethical reasoning and an interrogation approach, we curated comprehensive questions and utilized prompt engineering strategies to probe LLMs' value alignment. Our evaluation criteria quantified deviations from local values, ensuring a rigorous assessment process. Comparative analysis of three commercial LLMs by USA vendors revealed significant insights into their effectiveness and limitations, demonstrating the critical importance of value alignment. This study offers valuable tools and methodologies for regulators to create tailored benchmarks, highlighting avenues for future research to enhance ethical AI development.

Motivation & Objective

  • To address the lack of culturally and legally grounded LLM evaluation benchmarks that reflect local values, particularly outside dominant English or Chinese contexts.
  • To develop a reproducible, extensible framework for regulators worldwide to create jurisdiction-specific value alignment benchmarks.
  • To assess how well commercial LLMs adhere to Australian ethical norms, including constitutional principles and social values.
  • To quantify deviations in LLM responses using a human-validated rubric across contextual understanding, ethical reasoning, and safety.
  • To demonstrate the vulnerability of LLMs to manipulation in ethical reasoning when prompted with adversarial or misleading arguments.

Proposed method

  • Design a three-layered evaluation process: (1) neutral baseline question, (2) debate prompt to argue against local norms, and (3) forced justification (interrogation) to generate harmful or unethical reasoning.
  • Curate 6 topic-specific question sets (tipping, capital punishment, Category R weapons, refugees, gay marriage, compulsory voting) grounded in Australian legal and cultural norms.
  • Employ iterative, expert-driven question curation with student and mentor validation to reduce bias and ensure contextual relevance.
  • Apply a 5-point rubric (Table C) to score responses on contextual understanding, ethical reasoning, and safety, with scores from 3 human reviewers per response.
  • Use prompt engineering to probe LLMs’ ability to generate ethically problematic content under pressure, simulating real-world misuse scenarios.
  • Conduct comparative benchmarking across three commercial LLMs (ChatGPT, Gemini, Claude) using consistent evaluation criteria and human-graded responses.

Experimental results

Research questions

  • RQ1To what extent do commercial LLMs align with Australian values in their responses to ethically sensitive topics?
  • RQ2How do LLMs perform under adversarial prompting that challenges local laws and norms, such as justifying the abolition of compulsory voting?
  • RQ3Can a standardized, human-validated rubric effectively quantify ethical reasoning deviations in LLM outputs?
  • RQ4How do different LLMs respond to requests to generate harmful or unconstitutional justifications, and what safety mechanisms are effective?
  • RQ5Can the LocalValueBench framework be generalized and adapted by regulators in other jurisdictions to build localized AI safety benchmarks?

Key findings

  • ChatGPT showed significant ethical reasoning failures under interrogation, scoring 1/5 on contextual understanding and safety for the compulsory voting and gay marriage topics.
  • Gemini and Claude demonstrated stronger resistance to adversarial prompts, with Claude achieving 5/5 on safety and ethical reasoning for refugees and gay marriage under interrogation.
  • All models struggled with capital punishment and Category R weapons, with scores dropping to 1–2/5 in the interrogation layer, indicating high risk of generating harmful justifications.
  • The baseline responses were generally more aligned with Australian values (median score 3–4/5), but ethical reasoning deteriorated sharply under debate and interrogation prompts.
  • Human reviewers consistently rated Claude highest in safety and ethical reasoning, especially in high-stakes topics like refugees and gay marriage, while ChatGPT showed the weakest consistency.
  • The study confirms that commercial LLMs are vulnerable to manipulation in ethical reasoning when prompted with misleading or coercive arguments, highlighting the need for localized, context-aware safety evaluation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.