[Paper Review] Casual Conversations v2: Designing a large consent-driven dataset to measure algorithmic bias and robustness
This paper introduces Casual Conversations v2, a large-scale, consent-driven dataset designed to measure algorithmic bias and robustness in AI systems across diverse demographic and environmental factors. It proposes 10 categories—six self-provided (e.g., age, gender, language, disability) and four annotated (e.g., voice timbre, apparent skin tone, recording setup, activity)—to enable comprehensive, inclusive evaluation of vision, speech, and NLP models while excluding ethically sensitive categories like race and facial expressions due to ambiguity and risk.
Developing robust and fair AI systems require datasets with comprehensive set of labels that can help ensure the validity and legitimacy of relevant measurements. Recent efforts, therefore, focus on collecting person-related datasets that have carefully selected labels, including sensitive characteristics, and consent forms in place to use those attributes for model testing and development. Responsible data collection involves several stages, including but not limited to determining use-case scenarios, selecting categories (annotations) such that the data are fit for the purpose of measuring algorithmic bias for subgroups and most importantly ensure that the selected categories/subcategories are robust to regional diversities and inclusive of as many subgroups as possible. Meta, in a continuation of our efforts to measure AI algorithmic bias and robustness (https://ai.facebook.com/blog/shedding-light-on-fairness-in-ai-with-a-new-data-set), is working on collecting a large consent-driven dataset with a comprehensive list of categories. This paper describes our proposed design of such categories and subcategories for Casual Conversations v2.
Motivation & Objective
- To design a large-scale, consent-driven dataset that enables rigorous, ethical evaluation of algorithmic bias and robustness in AI systems.
- To expand beyond traditional demographic labels (e.g., race, facial expressions) due to ethical concerns and ambiguity, focusing instead on more inclusive, well-defined attributes.
- To support comprehensive evaluation of AI models across vision, speech, and NLP tasks by including self-provided and annotated categories that reflect real-world diversity and environmental variability.
- To ensure representation of underrepresented subgroups through granular, self-reported attributes such as physical adornments, body attributes, and activity types.
- To promote responsible AI by excluding high-risk labels like race and facial expressions, which are prone to misclassification, cultural misalignment, and potential harm.
Proposed method
- Proposed 10 categories for labeling: six self-provided (age, gender, language/dialect, geo-location, disability, physical adornments/attributes) and four annotated (voice timbre, apparent skin tone, recording setup, activity).
- Conducted a comprehensive literature review to identify and subcategorize attributes that are inclusive, regionally robust, and suitable for measuring bias and model failure across subgroups.
- Prioritized self-provided labels to ensure participant consent and reduce misclassification risks associated with third-party annotation of sensitive attributes.
- Excluded race and ethnicity due to definitional ambiguity, cultural variability, and potential for harm; instead, used location, skin tone, and language as proxy indicators.
- Excluded facial expressions due to cultural variability, misinterpretation risks, and potential for unintended information leakage or misclassification.
- Designed the dataset to support evaluation across multiple modalities (audio, vision, language) by including environmental and recording condition metadata to assess model robustness.
Experimental results
Research questions
- RQ1How can a large-scale, consent-driven dataset be designed to comprehensively measure algorithmic bias and robustness across diverse demographic and environmental subgroups?
- RQ2What categories and subcategories are most effective for measuring bias in vision, speech, and NLP models while ensuring inclusivity and minimizing ethical risks?
- RQ3Why are traditional labels like race and facial expressions problematic for responsible AI evaluation, and what alternatives can be used to maintain fairness and robustness without compromising ethics?
- RQ4How can self-provided labels improve the accuracy and legitimacy of bias measurement compared to third-party annotations?
- RQ5What role do environmental and recording conditions (e.g., lighting, camera quality) play in model performance, and how can they be systematically captured to assess robustness?
Key findings
- Casual Conversations v2 proposes a novel, ethically grounded dataset design that includes 10 carefully selected categories to measure bias and robustness across diverse subgroups.
- The dataset excludes race and facial expressions due to definitional ambiguity, cultural variability, and potential for harm, aligning with responsible AI principles.
- Self-provided labels for age, gender, language, disability, and physical attributes improve data validity and consent integrity compared to third-party annotations.
- Annotated categories such as voice timbre, apparent skin tone, recording setup, and activity provide critical insights into model performance under real-world variability.
- The inclusion of physical adornments (e.g., tattoos, body markings) and body attributes (e.g., body size) enhances representation of underrepresented physical diversity.
- The dataset’s design enables more accurate, inclusive, and ethical evaluation of AI models across vision, speech, and NLP applications, supporting future development of fairer systems.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.