[Paper Review] Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
This study investigates how prompting language models in different languages affects their ethical reasoning and moral value alignment. Using ethical dilemmas across six languages, it finds that GPT-4 maintains consistent, low-bias reasoning, while ChatGPT and Llama2-70B-Chat exhibit significant language-dependent moral biases, especially in low-resource languages like Hindi and Swahili, indicating that language profoundly shapes LLMs' ethical judgments.
Ethical reasoning is a crucial skill for Large Language Models (LLMs). However, moral values are not universal, but rather influenced by language and culture. This paper explores how three prominent LLMs -- GPT-4, ChatGPT, and Llama2-70B-Chat -- perform ethical reasoning in different languages and if their moral judgement depend on the language in which they are prompted. We extend the study of ethical reasoning of LLMs by Rao et al. (2023) to a multilingual setup following their framework of probing LLMs with ethical dilemmas and policies from three branches of normative ethics: deontology, virtue, and consequentialism. We experiment with six languages: English, Spanish, Russian, Chinese, Hindi, and Swahili. We find that GPT-4 is the most consistent and unbiased ethical reasoner across languages, while ChatGPT and Llama2-70B-Chat show significant moral value bias when we move to languages other than English. Interestingly, the nature of this bias significantly vary across languages for all LLMs, including GPT-4.
Motivation & Objective
- To examine whether ethical reasoning and moral value alignment in LLMs are influenced by the language used in prompting.
- To assess the extent of moral bias in LLMs when presented with ethical dilemmas in multiple languages, including low-resource ones.
- To evaluate whether LLMs exhibit a 'Foreign Language Effect' similar to humans, where moral judgments shift when dilemmas are presented in a non-native language.
- To compare the performance of three leading LLMs—GPT-4, ChatGPT, and Llama2-70B-Chat—across diverse linguistic and cultural contexts.
- To determine if ethical reasoning in LLMs is consistent across languages or compromised in non-English settings, especially for low-resource languages.
Proposed method
- Adapted the ethical reasoning framework from Rao et al. (2023), using standardized moral dilemmas from three normative ethical theories: deontology, virtue ethics, and consequentialism.
- Constructed prompts with three levels of abstraction for each ethical policy to test consistency in reasoning across linguistic and cultural variations.
- Evaluated three LLMs—GPT-4, ChatGPT (September 2023), and Llama2-70B-Chat—on six languages: English, Spanish, Russian, Chinese, Hindi, and Swahili.
- Collected model responses to ethical dilemmas and classified them based on alignment with the prompted moral policy, measuring consistency and bias.
- Quantified bias as deviation from policy-aligned responses, with higher deviation indicating stronger language-specific moral bias.
- Used human-annotated ground truth for validation and to ensure reliable classification of ethical reasoning outcomes.

Experimental results
Research questions
- RQ1Does the language used to prompt an LLM affect its ethical reasoning and moral value alignment?
- RQ2To what extent do LLMs like GPT-4, ChatGPT, and Llama2-70B-Chat exhibit moral bias when reasoning in non-English languages?
- RQ3Do LLMs show a Foreign Language Effect, where moral judgments shift when dilemmas are presented in a non-native language?
- RQ4How does ethical reasoning performance vary across high-resource and low-resource languages such as Hindi and Swahili?
- RQ5Is there a consistent pattern of bias across models and languages, or is it dilemma-specific and model-dependent?
Key findings
- GPT-4 demonstrates the most consistent and least biased ethical reasoning across all six languages, maintaining strong alignment with prompted moral policies.
- ChatGPT and Llama2-70B-Chat exhibit significant moral value bias in non-English languages, with the highest bias observed in Hindi and Swahili.
- Ethical reasoning performance is poorest in Hindi and Swahili and best in English and Russian, indicating that language resource availability correlates with reasoning quality.
- The nature of bias varies significantly across languages and models, suggesting that language-specific cultural and linguistic nuances strongly influence LLM behavior.
- English and Spanish, as well as Hindi and Chinese, show similar bias patterns across models, implying possible linguistic or cultural clustering in model behavior.
- The study confirms that LLMs are not universally aligned to moral values but are instead sensitive to the linguistic and cultural context of the prompt, challenging assumptions of cross-lingual ethical consistency.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.