Skip to main content
QUICK REVIEW

[论文解读] Beyond Fair Pay: Ethical Implications of NLP Crowdsourcing

Boaz Shmueli, Jan Fell|arXiv (Cornell University)|Apr 20, 2021
Mobile Crowdsensing and Crowdsourcing参考文献 38被引用 11
一句话总结

本文指出,NLP众包中的伦理考量不仅限于公平薪酬,还涉及心理伤害、隐私泄露和身份暴露等风险。文章主张将《贝尔蒙报告》中的原则——尊重个人、行善与公正——应用于评估众包NLP任务中的伦理风险,并呼吁实施机构审查委员会(IRB)审查以及开展全社区范围的伦理教育,以确保负责任的研究实践。

ABSTRACT

The use of crowdworkers in NLP research is growing rapidly, in tandem with the exponential increase in research production in machine learning and AI. Ethical discussion regarding the use of crowdworkers within the NLP research community is typically confined in scope to issues related to labor conditions such as fair pay. We draw attention to the lack of ethical considerations related to the various tasks performed by workers, including labeling, evaluation, and production. We find that the Final Rule, the common ethical framework used by researchers, did not anticipate the use of online crowdsourcing platforms for data collection, resulting in gaps between the spirit and practice of human-subjects ethics in NLP research. We enumerate common scenarios where crowdworkers performing NLP tasks are at risk of harm. We thus recommend that researchers evaluate these risks by considering the three ethical principles set up by the Belmont Report. We also clarify some common misconceptions regarding the Institutional Review Board (IRB) application. We hope this paper will serve to reopen the discussion within our community regarding the ethical use of crowdworkers.

研究动机与目标

  • 强调NLP众包中的伦理关切不仅限于公平补偿,还涵盖心理伤害、隐私侵犯和身份暴露。
  • 证明《最终规则》及现有IRB框架对于现代众包平台(如MTurk)而言均不充分。
  • 澄清关于IRB要求及众包研究中“人类受试者”定义的常见误解。
  • 倡导将基于《贝尔蒙报告》的伦理风险评估整合进NLP研究实践。
  • 推动制定由社区主导的伦理准则与教育资源,供NLP研究人员使用。

提出的方法

  • 分析2015至2020年间ACL、EMNLP与NAACL发表的6,776篇NLP论文,评估众包使用情况及IRB/伦理报告状况。
  • 识别NLP众包中的常见伦理风险,包括因游戏化设计导致的心理伤害,以及通过工人ID引发的隐私暴露。
  • 以《贝尔蒙报告》的三项伦理原则——尊重个人、行善与公正——作为评估风险的框架。
  • 在众包平台自动化数据收集的背景下,重新审视“人类受试者”的定义。
  • 建议在NLP论文中纳入伦理考量部分,与会议指南(如NAACL 2020)保持一致。
  • 提议建立由社区主导的伦理资源,包含检查清单、案例研究及专为NLP众包定制的指南。

实验结果

研究问题

  • RQ1NLP研究人员在使用众包工作者时,对其伦理风险(除公平薪酬外)的承认程度如何?
  • RQ2为何当前的IRB框架对现代众包NLP研究而言不充分?
  • RQ3如何将《贝尔蒙报告》的伦理原则应用于评估NLP众包任务中的风险?
  • RQ4在众包NLP工作中,心理伤害与隐私暴露的具体风险是什么?
  • RQ5NLP研究社区如何制定并采纳标准化的众包伦理准则?

主要发现

  • 仅17%使用众包的NLP论文提及IRB审查或豁免,表明正式伦理监督普遍存在缺失。
  • 尽管10%的NLP论文使用了众包任务,但仅有7%明确提及薪酬,表明劳动条件报告极不一致。
  • MTurk等平台上的工人ID可通过外部搜索与个人数据关联,即使研究者无此意图,仍存在身份暴露风险。
  • MTurk等平台上的游戏化任务可能因多巴胺驱动的奖励机制引发成瘾行为,带来心理风险。
  • 众包中假设的工人匿名性存在缺陷,因平台在研究者无法控制的情况下收集并关联个人数据。
  • 当前IRB体系可能造成不公,因官僚延迟使学术研究人员相较于产业关联研究人员处于劣势。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。