Skip to main content
QUICK REVIEW

[论文解读] NLPositionality: Characterizing Design Biases of Datasets and Models

Sebastin Santy, Jenny T. Liang|arXiv (Cornell University)|Jun 2, 2023
Hate Speech and Cyberbullying Detection被引用 4
一句话总结

NLPositionality 提出了一套框架,通过衡量来自全球多样化志愿者群体的标注与数据集标签或模型预测之间的对齐程度,来量化 NLP 数据集和模型中的设计偏见。研究发现,数据集和模型与西方、白人、受过大学教育且年龄较轻的人群最为对齐,而非二元性别个体和非英语母语者则系统性地被边缘化。

ABSTRACT

Design biases in NLP systems, such as performance differences for different populations, often stem from their creator's positionality, i.e., views and lived experiences shaped by identity and background. Despite the prevalence and risks of design biases, they are hard to quantify because researcher, system, and dataset positionality is often unobserved. We introduce NLPositionality, a framework for characterizing design biases and quantifying the positionality of NLP datasets and models. Our framework continuously collects annotations from a diverse pool of volunteer participants on LabintheWild, and statistically quantifies alignment with dataset labels and model predictions. We apply NLPositionality to existing datasets and models for two tasks -- social acceptability and hate speech detection. To date, we have collected 16,299 annotations in over a year for 600 instances from 1,096 annotators across 87 countries. We find that datasets and models align predominantly with Western, White, college-educated, and younger populations. Additionally, certain groups, such as non-binary people and non-native English speakers, are further marginalized by datasets and models as they rank least in alignment across all tasks. Finally, we draw from prior literature to discuss how researchers can examine their own positionality and that of their datasets and models, opening the door for more inclusive NLP systems.

研究动机与目标

  • 为解决 NLP 系统中因研究者、数据集和模型的立场性而产生的设计偏见缺乏量化工具的问题。
  • 开发一种可扩展、连续且包容的方法,以衡量 NLP 系统与多样化人群的对齐程度。
  • 揭示并描述现有数据集和模型中的系统性偏见,特别是针对社会可接受性与仇恨言论检测任务。
  • 提高 NLP 研究中对立场性的认识,并鼓励研究者审视自身及其系统的位置背景。

提出的方法

  • 利用 LabintheWild 平台,从 87 个国家招募 1,096 名多样化标注员,持续收集标注数据。
  • 使用标准化的李克特量表评分来收集社会可接受性和毒性标注,与原始数据集或模型标签保持一致。
  • 通过相关性和一致性度量,统计量化标注员人口统计学特征与数据集/模型输出之间的对齐程度。
  • 在不需重新训练或访问模型的情况下,将该框架后置应用于现有数据集和模型(包括 GPT-4)。
  • 使用人口统计学自报数据(如国籍、教育程度、年龄、性别认同)分析不同身份群体的立场性影响。
  • 优先考虑参与者的动机和学习意愿,而非金钱报酬,以提升数据质量,并实现长期可持续的数据收集。

实验结果

研究问题

  • RQ1NLP 数据集和模型在多大程度上与来自全球多样化地区的标注员的人口统计背景对齐?
  • RQ2NLP 系统中的设计偏见在多大程度上反映了其创建者和原始标注员的立场性?
  • RQ3哪些人口群体与数据集标签和模型预测的对齐程度最低,表明其系统性地被边缘化?
  • RQ4连续的、由志愿者驱动的标注框架能否有效捕捉 NLP 系统中不断演变的立场性?
  • RQ5像 GPT-4 这类模型的立场性与微调模型及数据集相比如何?

主要发现

  • 共从 87 个国家的 1,096 名标注员处收集了 16,299 条标注,平均每天收集 38 条,持续时间超过一年。
  • 数据集和模型与西方、白人、受过大学教育且年龄较轻的人群对齐程度最高——这正是 WEIRD(西方、受教育、工业化、富裕、民主)人口特征的体现。
  • 在所有任务中,非二元性别个体和非英语母语者与数据集标签及模型预测的对齐程度最低,表明其系统性地被边缘化。
  • 原始数据集标注员与其自身标签之间表现出强烈对齐,凸显了创建者立场性在塑造数据集偏见中的作用。
  • 该框架成功识别出微调模型和通用大语言模型(如 GPT-4)中的立场性模式,揭示其与特权人口群体的一致性对齐。
  • 本研究证明,在 LabintheWild 等平台上进行持续的、以动机驱动的标注,可产生比付费众包更高质量的数据,并实现对设计偏见的长期监测。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。