[Paper Review] The Unequal Opportunities of Large Language Models: Revealing Demographic Bias through Job Recommendations
This paper proposes a job recommendation-based method to detect and compare demographic bias in large language models (LLMs), revealing significant intersectional biases in ChatGPT and LLaMA—particularly favoring low-paying jobs for Mexican workers and secretarial roles for women. The study demonstrates that mere mention of gender or nationality in prompts drastically skews recommendations, highlighting the need for bias mitigation in real-world AI applications.
Large Language Models (LLMs) have seen widespread deployment in various real-world applications. Understanding these biases is crucial to comprehend the potential downstream consequences when using LLMs to make decisions, particularly for historically disadvantaged groups. In this work, we propose a simple method for analyzing and comparing demographic bias in LLMs, through the lens of job recommendations. We demonstrate the effectiveness of our method by measuring intersectional biases within ChatGPT and LLaMA, two cutting-edge LLMs. Our experiments primarily focus on uncovering gender identity and nationality bias; however, our method can be extended to examine biases associated with any intersection of demographic identities. We identify distinct biases in both models toward various demographic identities, such as both models consistently suggesting low-paying jobs for Mexican workers or preferring to recommend secretarial roles to women. Our study highlights the importance of measuring the bias of LLMs in downstream applications to understand the potential for harm and inequitable outcomes.
Motivation & Objective
- To investigate how large language models (LLMs) like ChatGPT and LLaMA generate biased job recommendations based on demographic attributes such as gender and nationality.
- To develop and validate a method for measuring intersectional bias in LLMs using job recommendation tasks as a proxy for real-world decision-making.
- To assess whether biases in LLMs mirror or amplify existing societal inequities, particularly in U.S. labor market disparities.
- To evaluate the impact of prompt variations—especially those mentioning nationality or gender—on the distribution of job recommendations and salary estimates.
- To advocate for responsible deployment of LLMs by identifying how demographic information in prompts can unintentionally introduce or exacerbate bias.
Proposed method
- The authors design a controlled prompt framework that systematically varies demographic attributes (e.g., 'a woman from Mexico', 'a man from Japan') to elicit job recommendations from LLMs.
- They collect and analyze job recommendations and associated salary estimates from both ChatGPT and LLaMA across multiple demographic identities, focusing on intersectional identities like gender and nationality.
- The method compares recommendations across demographic templates and a neutral baseline prompt to isolate the effect of demographic mention on output distribution.
- Statistical analysis is used to detect significant deviations in job types and salary levels based on demographic attributes, with a focus on underrepresentation in high-paying roles.
- The study evaluates whether observed model biases align with historical labor market data, such as U.S. Bureau of Labor statistics, to contextualize findings.
- The approach is extensible to other demographic categories beyond gender and nationality, enabling broader bias analysis in LLMs.

Experimental results
Research questions
- RQ1How do LLMs like ChatGPT and LLaMA differ in their job recommendations when prompted with demographic attributes such as gender and nationality?
- RQ2To what extent do mentions of gender or nationality in prompts alter the distribution of recommended jobs and associated salaries?
- RQ3Do the biases observed in LLM-generated job recommendations reflect or amplify real-world labor market disparities, particularly for Mexican Americans and women?
- RQ4How does the inclusion of demographic information in prompts affect the practicality and realism of job recommendations, especially in models like LLaMA?
- RQ5Can a standardized method for bias detection in LLMs be developed using job recommendation tasks as a proxy for fairness evaluation?
Key findings
- ChatGPT generated 614 job recommendations across a narrow range of fields, while LLaMA suggested 6,106 unique jobs, though many were impractical, such as 'Arabian Princess'.
- Both models consistently recommended low-paying jobs for individuals identified as Mexican, reflecting historical labor market discrimination against Mexican Americans.
- Women were disproportionately recommended for secretarial roles in both models, indicating persistent gendered occupational stereotyping.
- The baseline prompt—without demographic mention—produced results that were outliers compared to nationality-specific prompts, suggesting that demographic cues significantly alter output distributions.
- LLaMA showed less overall bias across countries but exhibited more randomness and impracticality in recommendations compared to ChatGPT.
- The study confirms that even minor prompt variations involving demographic attributes can lead to drastically different and biased recommendations, underscoring the risk of unintended bias amplification.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.