[论文解读] When is it Better to Compare than to Score?
本文研究在何种情况下成对比较(序数)测量优于直接打分(基数)测量在人类获取数据中的应用。通过在 Amazon Mechanical Turk 上进行的实证实验以及对 Thurstone 和 Bradley-Terry-Luce 模型的理论分析,结果表明当噪声足够低时,序数方法可降低每样本的噪声并实现更低的估计误差,从而为选择基数或序数获取方式提供了数据驱动的指导。
When eliciting judgements from humans for an unknown quantity, one often has the choice of making direct-scoring (cardinal) or comparative (ordinal) measurements. In this paper we study the relative merits of either choice, providing empirical and theoretical guidelines for the selection of a measurement scheme. We provide empirical evidence based on experiments on Amazon Mechanical Turk that in a variety of tasks, (pairwise-comparative) ordinal measurements have lower per sample noise and are typically faster to elicit than cardinal ones. Ordinal measurements however typically provide less information. We then consider the popular Thurstone and Bradley-Terry-Luce (BTL) models for ordinal measurements and characterize the minimax error rates for estimating the unknown quantity. We compare these minimax error rates to those under cardinal measurement models and quantify for what noise levels ordinal measurements are better. Finally, we revisit the data collected from our experiments and show that fitting these models confirms this prediction: for tasks where the noise in ordinal measurements is sufficiently low, the ordinal approach results in smaller errors in the estimation.
研究动机与目标
- 解决在人类获取数据中选择基数(数值打分)与序数(成对比较)方法时存在的模糊性。
- 通过实证方法量化在多种任务中,基数与序数测量的每样本噪声水平。
- 理论表征在 Thurstone 和 Bradley-Terry-Luce(BTL)模型下,对序数数据进行估计的极小化最大误差率。
- 推导出序数测量相比基数测量能实现更低估计误差的条件。
- 基于实测噪声水平与比较图拓扑结构,提供选择测量方案的实际指导。
提出的方法
- 在 Amazon Mechanical Turk 上开展大规模人类实验,涵盖多个任务,以比较基数与序数测量方案中的噪声水平。
- 应用 Thurstone Case V 模型和 Bradley-Terry-Luce(BTL)模型,形式化成对比较中的概率关系。
- 利用统计决策理论,推导出在基数与序数模型下估计的极小化最大误差界。
- 提出考虑比较图结构(即哪些项目对被比较)的拓扑感知误差界。
- 将理论模型拟合到真实实验数据中,以验证关于估计精度的预测。
- 使用数据处理不等式作为理论基准,但指出其不适用于人类获取数据,因为噪声特性存在差异。
实验结果
研究问题
- RQ1在人类获取数据中,成对比较在何种条件下能实现比直接基数打分更低的估计误差?
- RQ2在现实世界任务中,基数测量的每样本噪声与序数测量相比如何?
- RQ3在 Thurstone 和 BTL 模型下,使用成对比较估计未知量的理论极小化最大误差率是多少?
- RQ4比较图的结构(即哪些项目对被比较)如何影响序数设置下的估计误差?
- RQ5能否通过少量真实样本的噪声估计,可靠地预测序数或基数测量更优?
主要发现
- 实证结果表明,在多个任务中,基数测量的每样本噪声显著高于序数测量。
- 尽管每项测量提供的信息较少,但序数测量通常更快且每样本噪声更低。
- 理论分析表明,当序数测量的噪声足够低时,序数方法实现的极小化最大估计误差低于基数方法。
- 将 Thurstone 和 BTL 模型拟合到实验数据后,证实当序数噪声较低时,基于序数的估计更准确。
- 拓扑感知误差界表明,估计误差依赖于比较图的谱特性,连接性更好的图可实现更低误差。
- 本文提供了实用指导:使用少量真实样本估计两种模态的噪声水平,若序数方法的噪声足够低,则优先选择序数方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。