[Paper Review] About Face: A Survey of Facial Recognition Evaluation
This survey analyzes 100+ face datasets (1976–2019) totaling 145 million images, evaluates how benchmarks and demographics have evolved, and argues for explicit contextual reporting to align evaluation with real-world deployments.
We survey over 100 face datasets constructed between 1976 to 2019 of 145 million images of over 17 million subjects from a range of sources, demographics and conditions. Our historical survey reveals that these datasets are contextually informed, shaped by changes in political motivations, technological capability and current norms. We discuss how such influences mask specific practices (some of which may actually be harmful or otherwise problematic) and make a case for the explicit communication of such details in order to establish a more grounded understanding of the technology's function in the real world.
Motivation & Objective
- Capture how facial recognition evaluation has evolved across four historical periods and how dataset design shapes model performance
- Assess data sources, consent, privacy, and demographic representation in benchmarks
- Highlight gaps between benchmark performance and real-world outcomes and advocate for contextual reporting
- Propose improvements to evaluation reporting and governance to better reflect deployment contexts
Proposed method
- Historical, period-based analysis of 133 datasets (1976–2019) totaling 145,143,610 images of 17,733,157 individuals
- Categorization of data sources (photography sessions, web-sourced, surveillance), consent practices, and demographic reporting
- Analysis of evaluation metrics (FMR, FNMR, accuracy) and how threshold selection affects reported performance
- Cross-era synthesis of task types (detection, verification, identification, analysis) and their corresponding benchmarks
- Assessment of governance, auditing (e.g., NIST FVRT), and the need for holistic, deployment-aware evaluations
- Discussion of ethical risks, privacy concerns, and the potential for misuse in benchmarks and marketing
Experimental results
Research questions
- RQ1How have facial recognition benchmarks and data sources evolved from 1976 to 2019?
- RQ2What are the primary factors driving evaluation practices, including demographics, consent, and reporting norms?
- RQ3Why do benchmark results often diverge from real-world performance, and how can evaluations better reflect deployment contexts?
- RQ4What governance, auditing, and reporting improvements are needed to make evaluations more holistic and ethically responsible?
Key findings
- The survey covers 133 datasets (1976–2019) with 145,143,610 images of 17,733,157 subjects.
- Dataset releases show four periods with distinct trends in size, scope, and tasks, culminating in the deep learning era after 2014.
- Real-world deployment failures and biases (e.g., demographic disparities) are not always captured by benchmark performance.
- Data sources shifted from controlled photography to web-sourced and surveillance data, raising consent and privacy concerns.
- Demographic representation is uneven, with Western biases emerging in online datasets and problematic labeling practices in some datasets.
- Evaluation metrics (FMR, FNMR, accuracy) can be manipulated via thresholds; holistic auditing and context-aware reporting are recommended.
- NIST FVRT demonstrates the value of dual-mode evaluation (quantitative performance and qualitative usability) for deployment readiness.
- The paper advocates for explicit communication of dataset construction, consent, provenance, and intended use cases to ground evaluation in real-world function.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.