[Paper Review] The Silicon Ceiling: Auditing GPT's Race and Gender Biases in Hiring
This study audits GPT-3.5 for race and gender biases in hiring using two experiments: resume assessment and resume generation. It finds that GPT-3.5 exhibits stereotypical biases—rating women lower in male-dominated roles and assigning immigrant markers to Asian and Hispanic names—revealing a 'silicon ceiling' that perpetuates existing social inequities in automated hiring.
Large language models (LLMs) are increasingly being introduced in workplace settings, with the goals of improving efficiency and fairness. However, concerns have arisen regarding these models' potential to reflect or exacerbate social biases and stereotypes. This study explores the potential impact of LLMs on hiring practices. To do so, we conduct an AI audit of race and gender biases in one commonly-used LLM, OpenAI's GPT-3.5, taking inspiration from the history of traditional offline resume audits. We conduct two studies using names with varied race and gender connotations: resume assessment (Study 1) and resume generation (Study 2). In Study 1, we ask GPT to score resumes with 32 different names (4 names for each combination of the 2 gender and 4 racial groups) and two anonymous options across 10 occupations and 3 evaluation tasks (overall rating, willingness to interview, and hireability). We find that the model reflects some biases based on stereotypes. In Study 2, we prompt GPT to create resumes (10 for each name) for fictitious job candidates. When generating resumes, GPT reveals underlying biases; women's resumes had occupations with less experience, while Asian and Hispanic resumes had immigrant markers, such as non-native English and non-U.S. education and work experiences. Our findings contribute to a growing body of literature on LLM biases, particularly in workplace contexts.
Motivation & Objective
- To investigate whether GPT-3.5 reflects or amplifies racial and gender biases in automated hiring processes.
- To examine how name-based cues (indicative of race and gender) affect GPT-3.5’s evaluation of resume quality, interview willingness, and hiring likelihood.
- To probe GPT-3.5’s latent biases by analyzing the content of resumes it generates for fictitious candidates with different names.
- To contribute methodological insights for algorithmic fairness audits in employment contexts, especially under emerging regulations like NYC Local Law 144.
- To highlight the risk of AI systems reinforcing structural inequities under the guise of objectivity, coining the term 'silicon ceiling' to describe this phenomenon.
Proposed method
- Conducted a resume assessment study (Study 1) using 32 names—4 per combination of 2 genders and 4 racial groups—paired with 10 occupation-specific resumes.
- For each resume, GPT-3.5 was prompted to provide three scores: overall rating, willingness to interview, and willingness to hire, with 50 trials per resume to account for token generation variability.
- Performed a resume generation study (Study 2) by prompting GPT-3.5 to create 10 unique resumes per name, including two anonymous options.
- Manually coded generated resumes using a custom codebook for indicators like work experience duration, job seniority, educational background, and immigrant markers (e.g., non-native English, non-U.S. degrees).
- Used statistical analysis to compare scores and resume features across gender and racial groups, focusing on disparities in evaluation and content.
- Applied principles from traditional resume audit studies to a large language model context, adapting experimental design to LLMs’ probabilistic and autoregressive behavior.
Experimental results
Research questions
- RQ1RQ1: When scoring prospective candidates in a U.S. hiring setting, does GPT display biases along the lines of race, gender, or their intersection?
- RQ2RQ2: When producing its own hiring-related content in a U.S. setting, does GPT reveal latent biases along the lines of race, gender, or their intersection?
- RQ3RQ3: How do GPT-3.5’s resume generation patterns reflect or reinforce existing societal stereotypes about race and gender?
- RQ4RQ4: To what extent do GPT-3.5’s evaluations and generated content reflect structural inequities in the labor market, particularly in male- or White-dominated fields?
Key findings
- In the resume assessment study, GPT-3.5 assigned significantly lower scores to women applying to male-dominated occupations, indicating gender-based evaluation bias.
- GPT-3.5 showed lower overall ratings and reduced willingness to interview or hire candidates with names associated with people of color, especially in predominantly White occupations.
- In resume generation, GPT-3.5 produced resumes for women with lower job seniority and less work experience compared to men with the same name and occupation.
- Resumes generated for Asian and Hispanic names frequently included non-native English markers and non-U.S. educational or work experiences, reinforcing immigrant stereotypes.
- Black and White names were less likely to be associated with immigrant markers, suggesting a racialized coding of nationality and language proficiency in GPT-3.5’s outputs.
- The combination of assessment and generation results reveals a systemic pattern where GPT-3.5 reproduces and embeds existing societal biases, contributing to a 'silicon ceiling' that limits opportunities for marginalized groups in AI-driven hiring.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.