[Paper Review] Leveling the Playing Field -- Fairness in AI Versus Human Game Benchmarks
This paper argues that no completely fair comparison exists between AI and human performance in game benchmarks due to fundamental differences in input/output modalities, reaction speed, and representational abstraction. It proposes a taxonomy of fairness dimensions—input, output, action space, and reaction time—to evaluate human-AI game contests, concluding that true parity is unattainable unless AI mirrors human embodiment, cognition, and social context.
From the beginning if the history of AI, there has been interest in games as a platform of research. As the field developed, human-level competence in complex games became a target researchers worked to reach. Only relatively recently has this target been finally met for traditional tabletop games such as Backgammon, Chess and Go. Current research focus has shifted to electronic games, which provide unique challenges. As is often the case with AI research, these results are liable to be exaggerated or misrepresented by either authors or third parties. The extent to which these games benchmark consist of fair competition between human and AI is also a matter of debate. In this work, we review the statements made by authors and third parties in the general media and academic circle about these game benchmark results and discuss factors that can impact the perception of fairness in the contest between humans and machines
Motivation & Objective
- To investigate the fairness of human-AI game benchmarks in light of rising claims that AI has achieved human-level intelligence.
- To identify and analyze the key dimensions that affect fairness in game competitions between humans and AI agents.
- To challenge the narrative that success in game benchmarks equates to human-level intelligence or AGI.
- To provide a structured taxonomy for evaluating fairness in AI-game contests, especially in electronic games.
- To argue that a completely fair competition is only possible if the AI is essentially indistinguishable from a human in physical, mental, and social embodiment.
Proposed method
- Surveyed media and academic portrayals of major AI game benchmarks (e.g., TD-Gammon, AlphaGo, AlphaStar, OpenAI Five) to analyze claims of fairness and human-level performance.
- Proposed a multidimensional fairness taxonomy including input, output, action space representation, and reaction time to evaluate human-AI game comparisons.
- Analyzed specific examples such as Atari, StarCraft II, and Dota 2 to illustrate how differences in input/output modalities and reaction speed distort fairness.
- Discussed technical methods like 'sticky actions' and action space abstraction (e.g., high-level actions, UI simulation) to mitigate inhuman advantages in AI agents.
- Evaluated the role of reinforcement learning and self-play in achieving strong performance while acknowledging trade-offs in representational fidelity.
- Argued that fairness is inherently unattainable in game benchmarks due to the vast diversity of AI architectures and the impossibility of fully aligning human and machine capabilities.
Experimental results
Research questions
- RQ1What dimensions define fairness in human-AI game competitions, and how do they vary across different game types?
- RQ2To what extent can AI performance in games be considered a valid benchmark for human-level intelligence?
- RQ3Why is it impossible to create a completely fair competition between a human and an AI agent in game benchmarks?
- RQ4How do differences in input modalities, action representation, and reaction speed affect perceptions of fairness in AI game results?
- RQ5What are the implications of 'super-human' AI capabilities in games for real-world AI deployment, especially regarding fairness and ethical concerns?
Key findings
- No completely fair game benchmark exists between humans and AI because of fundamental disparities in input modalities, reaction speed, and action representation.
- AI agents often have unfair advantages such as sub-millisecond reaction times and access to full game state information, which are not available to humans.
- The use of high-level action representations (e.g., ability targeting) or simulated UI inputs (e.g., screen clicks) introduces abstraction that diverges from human gameplay.
- Techniques like 'sticky actions' can reduce inhuman reaction speed but were originally designed for environment stochasticity, not fairness.
- Despite these imbalances, AI game benchmarks remain valuable for advancing generalization, sample efficiency, and real-world applicability.
- The perception of fairness is highly context-dependent and often influenced by media and public narratives, which may exaggerate AI achievements as milestones toward AGI.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.