Skip to main content
QUICK REVIEW

[论文解读] Distinguishing Question Subjectivity from Difficulty for Improved Crowdsourcing

Yuan Jin, Mark Carman|arXiv (Cornell University)|Feb 12, 2018
Mobile Crowdsensing and Crowdsourcing参考文献 21被引用 3
一句话总结

本文提出了 SDR(主观性与难度响应)模型,这是一种概率框架,显式建模问题难度,通过潜在工作者偏好因子隐式捕捉主观性。通过将主观性差异与客观难度区分开来,该模型在多种众包数据集上提升了答案质量预测和工作者响应预测的性能,同时提供了与人类验证一致的问题难度与主观性排序。

ABSTRACT

The questions in a crowdsourcing task typically exhibit varying degrees of difficulty and subjectivity. Their joint effects give rise to the variation in responses to the same question by different crowd-workers. This variation is low when the question is easy to answer and objective, and high when it is difficult and subjective. Unfortunately, current quality control methods for crowdsourcing consider only the question difficulty to account for the variation. As a result,these methods cannot distinguish workers personal preferences for different correct answers of a partially subjective question from their ability/expertise to avoid objectively wrong answers for that question. To address this issue, we present a probabilistic model which (i) explicitly encodes question difficulty as a model parameter and (ii) implicitly encodes question subjectivity via latent preference factors for crowd-workers. We show that question subjectivity induces grouping of crowd-workers, revealed through clustering of their latent preferences. Moreover, we develop a quantitative measure of the subjectivity of a question. Experiments show that our model(1) improves the performance of both quality control for crowd-sourced answers and next answer prediction for crowd-workers,and (2) can potentially provide coherent rankings of questions in terms of their difficulty and subjectivity, so that task providers can refine their designs of the crowdsourcing tasks, e.g. by removing highly subjective questions or inappropriately difficult questions.

研究动机与目标

  • 解决现有众包质量控制方法的局限性,即混淆了由主观性引起的工作员分歧与由难度导致的错误。
  • 不将主观性视为固定属性,而是通过聚类工作者响应模式的潜在偏好因子来建模主观性。
  • 开发统一的概率框架,分别估计难度与主观性,从而更好地预测正确答案与个体工作者响应。
  • 提供与人类评估一致的主观性定量度量。
  • 通过帮助任务发布者识别并优化高度主观或过于困难的问题,支持任务设计。

提出的方法

  • SDR 模型使用带有潜在变量的概率图模型,表示每个问题的工作者专业知识和偏好因子。
  • 问题难度被建模为直接影响正确响应概率的直接参数。
  • 主观性通过潜在偏好因子隐式编码,这些因子基于工作者对部分主观问题的响应模式,将工作者聚类为不同组别。
  • 该模型采用贝叶斯推断方法,同时估计工作者专业知识、偏好因子以及问题难度与主观性参数。
  • 它利用多选项问题中的答案相关性结构,以提高估计精度与群体检测能力。
  • 提出一种新颖的一致性评估方法,将模型估计结果与人类标注的难度与主观性排序进行比较,验证模型的可解释性。

实验结果

研究问题

  • RQ1概率模型能否有效区分众包任务中由问题主观性引起的工作员分歧与由问题难度引起的工作员分歧?
  • RQ2如何通过潜在工作者偏好因子隐式建模主观性?这种建模是否能导致工作者的可检测聚类?
  • RQ3与现有基线相比,SDR 模型在预测正确答案与个体工作者响应方面的准确性提升程度如何?
  • RQ4模型对难度与主观性的估计能否与人类标注的这些属性排序产生有意义的相关性?
  • RQ5该模型能否通过识别并按难度与主观性水平对问题进行排序,支持任务设计?

主要发现

  • 在七个部分主观数据集上,SDR 模型在预测未见工作者响应方面显著优于五个基线模型,平均准确率提升 1.5%–2.5%。
  • Nemenyi 事后检验确认,在 α=0.10 水平下,SDR 与所有其他模型(CDS、GLAD、DS、MdWC)之间的性能差异具有统计显著性。
  • 在 Fashion 数据集上,该模型在预测未见响应方面实现了 0.7659 的平均准确率,优于次佳模型(CDS)的 0.7621。
  • SDR 估计的主观性与人类分配的排序之间存在强烈的负相关性(r ≈ -0.85),证实了其在主观性度量上的有效性。
  • 该模型在时尚判断任务中成功识别出三组不同的工作者群体,对应不同的偏好模式,揭示了由主观性引发的聚类现象。
  • 人类评估者对难度与主观性的排序与 SDR 的估计高度一致,相关系数分别为 r ≈ 0.82(难度)与 r ≈ 0.88(主观性)。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。