[论文解读] GPT detectors are biased against non-native English writers
GPT检测器将许多非母语英语的作文误分类为AI生成,而以英语为母语的写作则被正确识别;简单提示可以绕过检测器,提升教育与评估中的伦理风险。
The rapid adoption of generative language models has brought about substantial advancements in digital communication, while simultaneously raising concerns regarding the potential misuse of AI-generated content. Although numerous detection methods have been proposed to differentiate between AI and human-generated content, the fairness and robustness of these detectors remain underexplored. In this study, we evaluate the performance of several widely-used GPT detectors using writing samples from native and non-native English writers. Our findings reveal that these detectors consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified. Furthermore, we demonstrate that simple prompting strategies can not only mitigate this bias but also effectively bypass GPT detectors, suggesting that GPT detectors may unintentionally penalize writers with constrained linguistic expressions. Our results call for a broader conversation about the ethical implications of deploying ChatGPT content detectors and caution against their use in evaluative or educational settings, particularly when they may inadvertently penalize or exclude non-native English speakers from the global discourse. The published version of this study can be accessed at: www.cell.com/patterns/fulltext/S2666-3899(23)00130-7
研究动机与目标
- 评估公开可用的GPT检测器在母语英语与非母语英语写作样本上的公平性和鲁棒性。
- 在不同检测器中量化非母语作者的误报率和母语作者的漏报率。
- 研究语言增强或提示是否会影响检测器的表现。
- 检查检测器对困惑度的依赖是否导致对非母语作者的偏见。
- 为AI内容检测器的更安全、更加公平的使用提供建议。
提出的方法
- 对TOEFL作文(非母语作者)和美国八年级作文(母语作者)进行七种现成GPT检测器的评估。
- 计算跨检测器的误报率以及AI生成分类的一致性。
- 分析两组之间的困惑度差异并与检测结果相关联。
- 使用ChatGPT提示来提升或简化语言,并评估对误分类率和困惑度的影响。
- 测试二轮自我编辑提示以评估检测器绕过潜力。
- 通过使用ICLR 2023已接受论文的跨领域检查来补充分析,以评估母语与非母语作者之间困惑度的差异。
实验结果
研究问题
- RQ1在多个检测器中,GPT检测器是否对非母语英语写作的误报率高于对母语写作?
- RQ2语言增强或提示策略是否能降低检测器偏见,或反过来使检测器被绕过?
- RQ3困惑度是否是区分母语/非母语写作的AI生成文本的可靠独立信号?
- RQ4在学术写作场景(如会议摘要)中,检测器偏见如何体现,超出TOEFL/大学作文?
主要发现
- 检测器将超过一半的非母语TOEFL作文误分类为AI生成(平均误报率:61.22%)。
- 检测器对91篇TOEFL作文中的18篇作出AI生成的一致判断,而在91篇中有89篇被至少一个检测器标记。
- 通过ChatGPT用母语者般的词汇选择来提升非母语作文,误分类率从61.22%降至11.77%(1/91被一致判定为AI写作)。
- 相反,将母语大学作文简化为类似非母语写作的风格,误分类率上升至56.65%。
- 第二轮自我编辑提示可显著降低检测率(在某些情况下从最高100%降至高达13%),并提高困惑度,显示出对提示设计的易感性。
- 对ICLR 2023摘要的分析显示非母语作者在摘要中的困惑度较低,支持语言变异性与检测器偏见之间的联系。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。