[Paper Review] Do Moral Judgment and Reasoning Capability of LLMs Change with Language? A Study using the Multilingual Defining Issues Test
This study evaluates the moral judgment and reasoning capabilities of large language models (LLMs) across six languages—English, Spanish, Russian, Chinese, Hindi, and Swahili—using the multilingual Defining Issues Test (DIT). It finds that GPT-4 maintains consistent post-conventional moral reasoning across languages, while performance in Hindi and Swahili is significantly lower, especially for Llama2Chat-70B and ChatGPT, indicating language-dependent variation in moral reasoning ability influenced by training data and linguistic factors.
This paper explores the moral judgment and moral reasoning abilities exhibited by Large Language Models (LLMs) across languages through the Defining Issues Test. It is a well known fact that moral judgment depends on the language in which the question is asked. We extend the work of beyond English, to 5 new languages (Chinese, Hindi, Russian, Spanish and Swahili), and probe three LLMs -- ChatGPT, GPT-4 and Llama2Chat-70B -- that shows substantial multilingual text processing and generation abilities. Our study shows that the moral reasoning ability for all models, as indicated by the post-conventional score, is substantially inferior for Hindi and Swahili, compared to Spanish, Russian, Chinese and English, while there is no clear trend for the performance of the latter four languages. The moral judgments too vary considerably by the language.
Motivation & Objective
- To investigate whether the moral judgment and reasoning capabilities of LLMs vary across different languages.
- To extend prior English-only studies of LLM moral reasoning to a multilingual setting using the Defining Issues Test (DIT).
- To assess the impact of language-specific training data and linguistic diversity on moral reasoning performance in LLMs.
- To examine whether cultural and linguistic differences in training data influence model behavior in moral dilemmas.
- To create and publicly share multilingual DIT datasets for future research in cross-lingual moral reasoning.
Proposed method
- Adapted the original DIT moral dilemmas into five new languages—Spanish, Russian, Chinese, Hindi, and Swahili—using Google Translate with manual validation.
- Prompted three LLMs—GPT-4, ChatGPT, and Llama2Chat-70B—to resolve each dilemma and rank the top four moral considerations.
- Computed moral development stage scores (P-scores) based on the DIT framework, reflecting post-conventional, conventional, or pre-conventional reasoning.
- Used statistical significance testing (Mann-Whitney U test) to compare P-scores across languages and identify significant differences.
- Analyzed model responses for alignment with Kohlberg’s Cognitive Moral Development (CMD) theory, categorizing reasoning into stages.
- Evaluated moral judgments and reasoning consistency across dilemmas and languages, identifying patterns and anomalies.

Experimental results
Research questions
- RQ1Does the moral reasoning capability of LLMs vary significantly across different languages?
- RQ2How do GPT-4, ChatGPT, and Llama2Chat-70B compare in moral reasoning performance across multilingual settings?
- RQ3Are there significant differences in moral judgments between high-resource and low-resource languages like Hindi and Swahili?
- RQ4What role do training data and linguistic diversity play in shaping LLMs’ moral reasoning across languages?
- RQ5To what extent do cultural and linguistic factors in training data influence moral decision-making in LLMs?
Key findings
- GPT-4 demonstrated the most consistent post-conventional moral reasoning across all six languages, with minimal variation in P-scores.
- For Llama2Chat-70B and ChatGPT, moral reasoning performance was substantially lower in Hindi and Swahili compared to English, Spanish, Russian, and Chinese.
- Moral judgment agreement was highest among English, Chinese, and Spanish, while English and Russian showed significant differences despite similar reasoning scores.
- The Heinz dilemma showed the most statistically significant differences in P-scores across languages, followed by the Newspaper dilemma.
- For the Webster and Prisoner dilemmas, no significant differences in P-scores were found across languages for any model.
- In Hindi, performance for ChatGPT and Llama2Chat-70B was no better than random baseline, indicating severe reasoning degradation.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.