Skip to main content
QUICK REVIEW

[论文解读] Crowd & Prejudice: An Impossibility Theorem for Crowd Labelling without a Gold Standard

Nicolás Della Penna, Mark D. Reid|arXiv (Cornell University)|Apr 16, 2012
Auction Theory and Applications参考文献 18被引用 11
一句话总结

该论文表明,缺乏真实标签数据的众包标注算法在本质上存在缺陷,原因在于存在一种策略性均衡,即所有标注员均报告一种共享的、无信息量的‘偏见’,而非真实标签。核心结论是一个不可能性定理:在无真实标签数据的情况下,机制无法可靠地区分真实标注与偏见标注;但即使少量真实标签数据,也足以消除此类均衡,并使真实标注成为占优策略。

ABSTRACT

A common use of crowd sourcing is to obtain labels for a dataset. Several algorithms have been proposed to identify uninformative members of the crowd so that their labels can be disregarded and the cost of paying them avoided. One common motivation of these algorithms is to try and do without any initial set of trusted labeled data. We analyse this class of algorithms as mechanisms in a game-theoretic setting to understand the incentives they create for workers. We find an impossibility result that without any ground truth, and when workers have access to commonly shared 'prejudices' upon which they agree but are not informative of true labels, there is always equilibria where all agents report the prejudice. A small amount amount of gold standard data is found to be sufficient to rule out these equilibria.

研究动机与目标

  • 将无真实标签数据的众包标注算法建模为博弈论机制,以揭示其隐藏激励机制。
  • 阐明当标注员共享非信息性的共同偏见时,为何此类算法无法有效获取真实标注。
  • 证明即使少量真实标签数据,也能消除基于偏见的均衡,并恢复真实标注。
  • 建立机制可可靠识别知情标注员而无需事先真实标签的条件。

提出的方法

  • 将众包标注建模为博弈论机制,标注员为最大化自身收益而战略性选择标签。
  • 引入‘偏见’概念——一种标注员可协调的共享非信息性标注模式。
  • 分析无真实标签数据时的纳什均衡,表明所有标注员可能理性地选择偏见而非真实标注。
  • 利用真实标签数据识别偏离真实标签的个体,通过标签一致性将信任外推至其他标注员。
  • 通过标注数据点的重叠(或特征到标签的映射关系)将已验证的真实标签标注员的信任传播至其他个体。
  • 证明真实标签数据使真实标注成为占优策略,并消除基于偏见的均衡。

实验结果

研究问题

  • RQ1当标注员共享一种共同的、无信息量的偏见时,无真实标签数据的众包标注机制能否可靠识别真实标签?
  • RQ2在缺乏真实标签的情况下,会出现何种策略性均衡?这些均衡是否阻碍真实标签的报告?
  • RQ3访问少量真实标签数据如何破坏基于共享偏见的均衡?
  • RQ4在何种条件下,真实标签数据能确保真实标注成为知情标注员的占优策略?
  • RQ5利用真实标签数据的机制能否通过标签一致性模式识别出偏见标注员和真实标注员?

主要发现

  • 在无真实标签数据的情况下,总存在一种纳什均衡,即所有标注员报告一种共享的、无信息量的偏见,无论其真实知识如何。
  • 在这些均衡中,机制无法区分真实标注与偏见标注,因为其观察到的行为完全相同。
  • 即使存在少量真实标签数据,也能通过机制识别出偏离真实标签的个体,从而打破基于偏见的均衡。
  • 真实标签数据使机制能够识别真实标注员,并通过标签一致性传播信任,最终正确分类所有标注员。
  • 在真实标签数据存在的情况下,真实标注成为知情标注员的占优策略,消除了跟随偏见的激励。
  • 机制可通过数据点重叠或特征到标签映射关系的重叠,成功识别出真实标注员和偏见标注员。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。