Skip to main content
QUICK REVIEW

[论文解读] A Technical Survey on Statistical Modelling and Design Methods for Crowdsourcing Quality Control

Yuan Jin, Mark Carman|arXiv (Cornell University)|Dec 5, 2018
Mobile Crowdsensing and Crowdsourcing参考文献 127被引用 5
一句话总结

本文对众包质量控制中的统计建模与机制设计进行了统一的技术综述,弥合了两个传统上分离的研究领域。它提出了整合统计推断与激励机制的综合框架,提供了关于主观性感知建模、多类别/标签处理以及支付与游戏化设计实证验证的分类体系与开放性挑战。

ABSTRACT

Online crowdsourcing provides a scalable and inexpensive means to collect knowledge (e.g. labels) about various types of data items (e.g. text, audio, video). However, it is also known to result in large variance in the quality of recorded responses which often cannot be directly used for training machine learning systems. To resolve this issue, a lot of work has been conducted to control the response quality such that low-quality responses cannot adversely affect the performance of the machine learning systems. Such work is referred to as the quality control for crowdsourcing. Past quality control research can be divided into two major branches: quality control mechanism design and statistical models. The first branch focuses on designing measures, thresholds, interfaces and workflows for payment, gamification, question assignment and other mechanisms that influence workers' behaviour. The second branch focuses on developing statistical models to perform effective aggregation of responses to infer correct responses. The two branches are connected as statistical models (i) provide parameter estimates to support the measure and threshold calculation, and (ii) encode modelling assumptions used to derive (theoretical) performance guarantees for the mechanisms. There are surveys regarding each branch but they lack technical details about the other branch. Our survey is the first to bridge the two branches by providing technical details on how they work together under frameworks that systematically unify crowdsourcing aspects modelled by both of them to determine the response quality. We are also the first to provide taxonomies of quality control papers based on the proposed frameworks. Finally, we specify the current limitations and the corresponding future directions for the quality control research.

研究动机与目标

  • 弥合众包质量控制中统计建模与机制设计之间的差距,这两者历史上一直被孤立研究。
  • 提供一个统一的技术框架,系统性地整合建模与设计两方面,以改善响应质量的推断。
  • 基于统一框架,提出现有质量控制方法的详细分类体系,提升对文献的分类与理解。
  • 识别并阐明当前方法的关键局限,特别是关于主观性、多类别/标签处理以及激励机制实证验证方面的问题。
  • 概述未来研究方向,包括主观性建模、响应选项中语义关系的处理,以及开发基于实证的 gamification 方法论。

提出的方法

  • 本文提出一个统一框架,将统计模型(用于响应聚合与质量推断)与机制设计(用于激励、阈值与工作流)相连接。
  • 提出一种联合建模方法,其中统计模型估计工人可靠性、任务难度与主观性,而机制则利用这些估计值设定阈值与支付规则。
  • 该框架将工人专业度、任务难度与响应主观性作为潜在变量纳入概率模型,从而从噪声响应中推断真实标签。
  • 作者采用贝叶斯推断方法估计模型参数,核心方程基于工人准确率与任务特征建模响应似然。
  • 该综述包含对100余篇质量控制论文的分类,按其建模假设与设计机制分类,以厘清研究格局与方法论差距。
  • 提出矩阵分解技术,用于在高度多类别或多标签任务中建模响应选项之间的语义关系,降低模型复杂度并防止过拟合。

实验结果

研究问题

  • RQ1如何系统性地统一统计模型与机制设计,以提升众包质量控制?
  • RQ2在质量控制框架中,统计推断与激励机制之间的关键技术关联是什么?
  • RQ3如何对响应中的主观性进行定量建模,并将其整合进质量控制系统?
  • RQ4当前方法在处理高度多类别或多标签众包任务方面存在哪些局限?
  • RQ5如何对游戏化与支付机制进行实证验证与优化,以提升响应质量?

主要发现

  • 本文识别出一个关键研究空白:现有综述将统计建模与机制设计视为独立领域,尽管在真实系统中二者存在强烈依赖关系。
  • 研究证明,统计模型可提供关键的参数估计与理论性能保证,从而指导机制设计,如最优阈值与支付规则的设定。
  • 仅有三篇研究处理了高度多类别众包中的质量控制问题,且仅有一项研究(采用基于矩阵的建模)展示了从响应中重建语义关系的潜力。
  • 主观性感知模型仍处于初步阶段,仅有1项研究(Yuan et al. [31])提出基于不同工人群体预期正确答案的问题特定主观性度量。
  • 当前的游戏化设计缺乏实证验证,更多依赖轶事性或电子游戏启发的惯例,而非数据驱动的指导原则。
  • 本文强调,矩阵分解在大规模选项任务中对语义关系建模至关重要,可降低计算复杂度并防止过拟合。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。