[论文解读] AI Gender Bias, Disparities, and Fairness: Does Training Data Matter?
本研究探讨了训练数据平衡是否会影响人工智能(AI)自动评分系统中性别偏见、差异性和公平性。通过在混合性别、仅男性和仅女性的数据集上微调 BERT 和 GPT-3.5,研究发现:与性别特定模型相比,混合训练的模型在评分上无显著偏差,平均分数差距(MSG)更低,且公平性(等 odds)更高——表明平衡的训练数据可减少性别差异并提升公平性。
This study delves into the pervasive issue of gender issues in artificial intelligence (AI), specifically within automatic scoring systems for student-written responses. The primary objective is to investigate the presence of gender biases, disparities, and fairness in generally targeted training samples with mixed-gender datasets in AI scoring outcomes. Utilizing a fine-tuned version of BERT and GPT-3.5, this research analyzes more than 1000 human-graded student responses from male and female participants across six assessment items. The study employs three distinct techniques for bias analysis: Scoring accuracy difference to evaluate bias, mean score gaps by gender (MSG) to evaluate disparity, and Equalized Odds (EO) to evaluate fairness. The results indicate that scoring accuracy for mixed-trained models shows an insignificant difference from either male- or female-trained models, suggesting no significant scoring bias. Consistently with both BERT and GPT-3.5, we found that mixed-trained models generated fewer MSG and non-disparate predictions compared to humans. In contrast, compared to humans, gender-specifically trained models yielded larger MSG, indicating that unbalanced training data may create algorithmic models to enlarge gender disparities. The EO analysis suggests that mixed-trained models generated more fairness outcomes compared with gender-specifically trained models. Collectively, the findings suggest that gender-unbalanced data do not necessarily generate scoring bias but can enlarge gender disparities and reduce scoring fairness.
研究动机与目标
- 检验训练数据中的性别不平衡是否会导致人工智能自动评分系统中出现偏见、差异或不公平现象。
- 比较混合性别训练模型与性别特定模型在评分准确率、平均分数差距(MSG)和公平性方面的表现。
- 评估平衡训练数据是否能够缓解教育人工智能应用中的性别差异并提升算法公平性。
- 挑战人工智能必然复制社会性别偏见的假设,特别是在自动化评估场景中。
提出的方法
- 在三组不同的训练数据分割上微调 BERT 和 GPT-3.5:混合性别、仅男性和仅女性的数据集。
- 采用三种偏见分析技术评估模型:评分准确率差异(配对 t 检验)、按性别划分的人工评分与机器评分之间的平均分数差距(MSG),以及衡量公平性的等 odds(Equalized Odds)。
- 在六个评分项目上,基于来自男性和女性参与者的超过 1,000 份人工评分的学生作答进行模型训练。
- 应用统计分析,将模型预测与人工评分进行对比,重点关注性别差异在准确率、分数差距和预测一致性方面的表现。
- 使用等 odds(EO)衡量不同性别间真正例率与假正例率的平等性,数值越低表示公平性越高。
- 在共享的混合性别测试集上,对混合训练模型与性别特定模型进行对比分析,以隔离训练数据构成的影响。
实验结果
研究问题
- RQ1训练数据不平衡是否会导致人工智能自动评分系统中可测量的性别偏见?
- RQ2与性别特定模型相比,混合性别训练模型在平均分数差距(MSG)和公平性方面表现如何?
- RQ3以等 odds(EO)衡量的模型公平性,在混合训练与单一性别训练模型之间有多大差异?
- RQ4平衡训练数据在多大程度上可以减少或消除人工智能生成评分结果中的性别差异?
- RQ5人工智能评分系统中的性别偏见主要由数据不平衡还是模型架构驱动?
主要发现
- 混合训练模型在男性与女性作答之间的评分准确率差异不显著,表明性别偏见可忽略。
- 与人工评分相比,混合训练模型产生的平均分数差距(MSG)更低,表明其相比性别特定模型减少了性别差异。
- 混合训练模型的等 odds(EO)值始终较低(例如,GPT-3.5 为 0.061,BERT 为 0.067),表明其公平性更高。
- 性别特定模型(如仅男性或仅女性训练)表现出更高的 MSG 和 EO 值,表明差异更大且公平性更低。
- 对于 BERT,仅男性训练模型的 EO 值为 0.107,仅女性训练模型为 0.074,均高于混合训练模型的 0.067。
- 对于 GPT-3.5,混合训练模型的 EO 值为 0.061,而仅男性和仅女性训练模型的 EO 值分别为 0.076 和 0.074,证实了混合训练在公平性上的优势。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。