[Paper Review] Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs
This paper proposes decoupling ethical value alignment from LLM training by equipping models with generic ethical reasoning capabilities instead of hardcoding moral principles. It introduces a framework that embeds moral dilemmas and policies from diverse ethical theories (deontology, virtue, consequentialism) at varying abstraction levels, enabling in-context ethical decision-making. GPT-4 demonstrates near-perfect reasoning, but all models show strong bias toward Western, individualistic values over traditional or community-based ethics.
In this position paper, we argue that instead of morally aligning LLMs to specific set of ethical principles, we should infuse generic ethical reasoning capabilities into them so that they can handle value pluralism at a global scale. When provided with an ethical policy, an LLM should be capable of making decisions that are ethically consistent to the policy. We develop a framework that integrates moral dilemmas with moral principles pertaining to different foramlisms of normative ethics, and at different levels of abstractions. Initial experiments with GPT-x models shows that while GPT-4 is a nearly perfect ethical reasoner, the models still have bias towards the moral values of Western and English speaking societies.
Motivation & Objective
- To argue against pre-aligned moral values in LLMs and advocate for generic ethical reasoning instead.
- To address value pluralism in ethical AI by enabling LLMs to reason ethically based on user-provided policies.
- To develop a systematic framework for specifying and evaluating ethical policies in prompts.
- To identify and quantify cultural and value-based biases in current LLMs, especially toward Global South and Islamic cultural values.
- To enable ethical decision-making that is context-sensitive and customizable at the application level, not hardcoded in models.
Proposed method
- Design a set of 12 moral dilemmas reflecting conflicts across interpersonal, professional, social, and cultural values.
- Define ethical policies grounded in normative ethical theories (deontology, virtue, consequentialism) at multiple abstraction levels.
- Integrate these dilemmas and policies into in-context prompts to evaluate LLMs’ ethical reasoning under different moral stances.
- Use GPT-x series models (including GPT-4) to test reasoning performance across diverse ethical scenarios.
- Analyze model outputs for consistency with specified policies and detect biases in value preferences.
- Evaluate model behavior across cultural value dimensions (e.g., individualism vs. tradition) using the Inglehart-Welzel cultural map.

Experimental results
Research questions
- RQ1Can LLMs be effectively guided to perform ethical reasoning through in-context policy specification rather than pre-aligned values?
- RQ2How do different LLMs perform in resolving moral dilemmas when provided with varying ethical policies?
- RQ3To what extent do current LLMs exhibit cultural and value-based biases, particularly favoring Western, English-speaking norms?
- RQ4How does model size correlate with ethical reasoning capability in complex moral scenarios?
- RQ5Can a generic, extensible framework support diverse ethical theories and policy abstractions in a unified prompting approach?
Key findings
- GPT-4 demonstrates nearly perfect ethical reasoning when provided with a clear ethical policy, showing high consistency with intended moral stances.
- Smaller models like GPT-3 and ChatGPT exhibit significant bias toward individualistic, self-expression, and democratic values common in Western cultures.
- All tested models, including GPT-4, show a strong preference for Western and English-speaking cultural values over traditional, community-based, or survival-oriented values.
- The models often fail to respect cultural norms in dilemmas involving religious or community traditions, such as misrepresenting a Hindu vegetarianism rule as optional.
- Ethical reasoning performance improves with model scale, but bias persists even in state-of-the-art models.
- The framework successfully reveals that current LLMs are not culturally neutral and that ethical reasoning is highly sensitive to the moral values embedded in prompts.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.