[Paper Review] What You See Is What You Get? The Impact of Representation Criteria on Human Bias in Hiring
This study investigates how representation criteria—fixed proportional gender distribution in candidate slates—affect human gender bias in hiring decisions. Using controlled experiments on Amazon Mechanical Turk, it finds that balancing gender representation mitigates bias in professions with skewed world distributions, but not where persistent human preferences override representation, with decision-maker gender and task complexity further influencing outcomes.
Although systematic biases in decision-making are widely documented, the ways in which they emerge from different sources is less understood. We present a controlled experimental platform to study gender bias in hiring by decoupling the effect of world distribution (the gender breakdown of candidates in a specific profession) from bias in human decision-making. We explore the effectiveness of \ extit{representation criteria}, fixed proportional display of candidates, as an intervention strategy for mitigation of gender bias by conducting experiments measuring human decision-makers' rankings for who they would recommend as potential hires. Experiments across professions with varying gender proportions show that balancing gender representation in candidate slates can correct biases for some professions where the world distribution is skewed, although doing so has no impact on other professions where human persistent preferences are at play. We show that the gender of the decision-maker, complexity of the decision-making task and over- and under-representation of genders in the candidate slate can all impact the final decision. By decoupling sources of bias, we can better isolate strategies for bias mitigation in human-in-the-loop systems.
Motivation & Objective
- To isolate and measure the impact of human decision-making bias in hiring, independent of world distribution and algorithmic bias.
- To evaluate whether representation criteria—fixed proportional gender display in candidate slates—can mitigate human gender bias in hiring decisions.
- To understand how factors like decision-maker gender, task complexity, and over/under-representation influence hiring outcomes.
- To decompose sources of bias in hybrid human-in-the-loop hiring systems by decoupling world, algorithmic, and human contributions.
- To provide empirical evidence for designing fairer hiring systems by identifying when representation criteria are effective or insufficient.
Proposed method
- Conducted large-scale controlled experiments using Amazon Mechanical Turk to simulate hiring decisions across diverse professions.
- Generated candidate profiles with identical qualifications but randomized gender and pronouns, maintaining fixed gender proportions in each slate.
- Applied representation criteria by ensuring candidate slates had predefined gender distributions (e.g., 50-50, 3-97) to isolate the effect of representation.
- Measured human recommendations by asking participants to rank their top 4 candidates out of 8, simulating a recommendation task.
- Compared outcomes against baselines using real-world gender distributions and an AI model trained on word embeddings to assess bias without intervention.
- Analyzed data by decision-maker gender and profession to assess confounding effects on decision-making.
Experimental results
Research questions
- RQ1Does balancing gender representation in candidate slates mitigate human gender bias in hiring decisions?
- RQ2How does the effectiveness of representation criteria vary across professions with different underlying gender distributions?
- RQ3In professions where representation criteria fail to reduce bias, does over-representation of underrepresented genders improve outcomes?
- RQ4How do personal characteristics of decision-makers—particularly their gender—affect hiring recommendations?
- RQ5How do task complexity and candidate slate composition (e.g., over- or under-representation) influence human decision-making in hiring?
Key findings
- Representation criteria significantly reduce gender bias in professions where the world distribution is highly skewed, such as software engineering and law.
- In professions with persistent human preferences—like plumbing or construction—representation criteria fail to correct for bias, even when gender distribution in slates is balanced.
- The gender of the decision-maker influences outcomes: female decision-makers show less bias than male decision-makers in certain professions, though effects vary.
- Task complexity and the degree of under- or over-representation in candidate slates confound hiring decisions, with under-representation amplifying bias.
- Over-representation of women in slates (e.g., 50% female in a male-dominated field) does not consistently improve fairness outcomes, especially where deep-seated preferences persist.
- The study demonstrates that representation criteria are not a universal solution; their effectiveness depends on the interplay between world distribution, human preferences, and decision-maker characteristics.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.