Skip to main content
QUICK REVIEW

[论文解读] Statistical Inference for Incomplete Ranking Data: The Case of Rank-Dependent Coarsening

Mohsen Ahmadi Fahandar, Eyke Hüllermeier|arXiv (Cornell University)|Dec 4, 2017
Game Theory and Voting Systems参考文献 17被引用 7
一句话总结

本文提出了一种统计框架,用于在依赖排名的粗化条件下对排序数据进行排序,其中不完全排序源于将完整排序投影到随机排名子集上。研究表明,尽管该粗化过程引入了偏差,但在Plackett-Luce假设下,多种排序方法——包括Copeland和FAS——仍具有一致性,即随着样本量增大,能够恢复真实排序。

ABSTRACT

We consider the problem of statistical inference for ranking data, specifically rank aggregation, under the assumption that samples are incomplete in the sense of not comprising all choice alternatives. In contrast to most existing methods, we explicitly model the process of turning a full ranking into an incomplete one, which we call the coarsening process. To this end, we propose the concept of rank-dependent coarsening, which assumes that incomplete rankings are produced by projecting a full ranking to a random subset of ranks. For a concrete instantiation of our model, in which full rankings are drawn from a Plackett-Luce distribution and observations take the form of pairwise preferences, we study the performance of various rank aggregation methods. In addition to predictive accuracy in the finite sample setting, we address the theoretical question of consistency, by which we mean the ability to recover a target ranking when the sample size goes to infinity, despite a potential bias in the observations caused by the (unknown) coarsening.

研究动机与目标

  • 解决从不完全排序数据中学习的统计挑战,其中完整排序通过投影到随机排名子集而被粗化。
  • 显式建模粗化过程,与现有方法形成对比,后者忽略或假设粗化与排序无关。
  • 研究当粗化引入偏差时,排序方法是否仍能保持一致性(即恢复真实排序)。
  • 提供理论和实证证据,证明在依赖排名的粗化下的一致性,特别是针对Plackett-Luce分布的完整排序。
  • 扩展对排序情境中粗化数据与粗化机制的理解,与基于项目的边际化方法区分开来。

提出的方法

  • 引入依赖排名的粗化概念,即完整排序被投影到随机排名子集上,产生不完全排序。
  • 使用符号τ表示不完全排序,其中τ(k) = 0表示项目k缺失,τ(k) > 0表示其被排序。
  • 将不完全排序τ的全部一致完整排序(线性扩展)集合定义为E(τ),以支持概率推断。
  • 假设完整排序服从参数向量θ的Plackett-Luce分布,并推导出在粗化后观察到成对偏好a_i ≻ a_j的概率。
  • 使用粗化机制λ_{r,s}将完整排序投影到随机排名对r和s上,并计算观察到a_i ≻ a_j的诱导概率q_{i,j}。
  • 通过证明p_{i,j} - 1/2与q'_{i,j} - 1/2的符号一致,建立理论一致性,确保在存在偏差的情况下仍能保持顺序不变。

实验结果

研究问题

  • RQ1当观测因依赖排名的粗化而产生偏差时,排序方法是否仍能保持一致性,以恢复真实排序?
  • RQ2Plackett-Luce模型在依赖排名的粗化后是否仍能保持成对偏好中项目相对顺序的稳定性?
  • RQ3在该粗化模型下,常用排序聚合方法如Copeland和FAS是否仍具有一致性,即使估计值存在偏差?
  • RQ4依赖排名的粗化与标准基于项目的边际化在排序推断的统计影响上如何不同?
  • RQ5在样本量增加时,什么条件能确保估计的成对偏好p̂_{i,j}正确反映真实偏好顺序?

主要发现

  • 在依赖排名的粗化下,Plackett-Luce模型保持了成对偏好差异符号的一致性:当且仅当q'_{i,j} > 1/2时,有p_{i,j} > 1/2。
  • 引理2表明,若潜在分布偏好与a_i ≻ a_j一致的排序,则粗化后的概率q_{i,j}仍大于q_{j,i}。
  • 引理3证明,在Plackett-Luce下,粗化过程保持顺序不变,确保任意两个项目之间的相对偏好不会被颠倒。
  • 定理5建立Copeland排序方法在所提模型下的一致性,即随着N → ∞,其以概率趋近于1恢复真实排序。
  • 定理6确认FAS、FAS(R)和FAS(B)的一致性,表明这些方法对依赖排名的粗化具有鲁棒性。
  • 实验结果表明一致性可能扩展至BTL等其他方法,但其在一般模型下的推广尚缺乏正式证明。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。