[论文解读] Uncertainty Quantification of MLE for Entity Ranking with Covariates
本文提出了协变量辅助排名估计(Covariate-Assisted Ranking Estimation, CARE)模型,通过在潜在得分结构 $\alpha_i^* + \mathbf{x}_i^\top\bm{\beta}^*$ 中引入项目协变量,扩展了 Bradley-Terry-Luce (BTL) 模型。该研究建立了最大似然估计量(MLE)在 $\ell_2$ 和 $\ell_\infty$ 范数下的最优估计速率,并推导出渐近分布以实现不确定性量化,从而在最小样本复杂度下实现对协变量效应的统计推断。
This paper concerns with statistical estimation and inference for the ranking problems based on pairwise comparisons with additional covariate information such as the attributes of the compared items. Despite extensive studies, few prior literatures investigate this problem under the more realistic setting where covariate information exists. To tackle this issue, we propose a novel model, Covariate-Assisted Ranking Estimation (CARE) model, that extends the well-known Bradley-Terry-Luce (BTL) model, by incorporating the covariate information. Specifically, instead of assuming every compared item has a fixed latent score $\{θ_i^*\}_{i=1}^n$, we assume the underlying scores are given by $\{α_i^*+{x}_i^ opβ^*\}_{i=1}^n$, where $α_i^*$ and ${x}_i^ opβ^*$ represent latent baseline and covariate score of the $i$-th item, respectively. We impose natural identifiability conditions and derive the $\ell_{\infty}$- and $\ell_2$-optimal rates for the maximum likelihood estimator of $\{α_i^*\}_{i=1}^{n}$ and $β^*$ under a sparse comparison graph, using a novel `leave-one-out' technique (Chen et al., 2019) . To conduct statistical inferences, we further derive asymptotic distributions for the MLE of $\{α_i^*\}_{i=1}^n$ and $β^*$ with minimal sample complexity. This allows us to answer the question whether some covariates have any explanation power for latent scores and to threshold some sparse parameters to improve the ranking performance. We improve the approximation method used in (Gao et al., 2021) for the BLT model and generalize it to the CARE model. Moreover, we validate our theoretical results through large-scale numerical studies and an application to the mutual fund stock holding dataset.
研究动机与目标
- 解决在现实世界中普遍存在但缺乏统计推断方法的、纳入项目协变量的实体排序模型问题。
- 提出一种新模型——协变量辅助排名估计(CARE)模型,通过将潜在得分建模为基线分量与协变量驱动分量之和,扩展 Bradley-Terry-Luce (BTL) 模型。
- 在稀疏比较图下,建立 MLE 对内在得分 $\alpha_i^*$ 和协变量系数 $\bm{\beta}^*$ 的最优统计收敛速率。
- 推导 MLE 的渐近分布,以实现对协变量效应的不确定性量化与假设检验。
- 提供一个实用的推断框架,用于评估协变量是否能解释潜在得分的变化,并通过阈值化稀疏参数提升排名性能。
提出的方法
- 提出 CARE 模型:$\mathbb{P}(\text{item } i \text{ preferred over } j) = \frac{e^{\alpha_i^* + \mathbf{x}_i^\top\bm{\beta}^*}}{e^{\alpha_i^* + \mathbf{x}_i^\top\bm{\beta}^*} + e^{\alpha_j^* + \mathbf{x}_j^\top\bm{\beta}^*}}$,其中 $\alpha_i^*$ 为基线分量,$\mathbf{x}_i^\top\bm{\beta}^*$ 为协变量效应分量。
- 使用带可识别性约束的约束最大似然估计量(MLE)$\widehat{\bm{\theta}}_M = (\widehat{\bm{\alpha}}_M, \widehat{\bm{\beta}}_M)$,以确保估计的唯一性。
- 应用一种新颖的“留一法”技术(Chen et al., 2019),在稀疏 Erdős-Rényi 比较图下推导 $\widehat{\bm{\alpha}}_M$ 和 $\widehat{\bm{\beta}}_M$ 的 $\ell_2$ 与 $\ell_\infty$ 收敛速率。
- 推导 $\bm{\beta}^*$ 和 $\alpha_i^*$ 的 MLE 的渐近正态性,从而支持对协变量效应的置信区间构造与假设检验。
- 改进 Gao et al. (2021) 对 BLT 模型的近似方法,并将其推广至 CARE 模型,以提升高维设置下的推断精度。
- 采用对 $\bm{\alpha}$ 进行 $\ell_2$ 正则化的投影梯度下降算法,以确保数值优化过程的收敛性与稳定性。
实验结果
研究问题
- RQ1我们能否设计一种在配对比较排名中有效结合项目协变量并具备可证明推断保证的统计高效模型?
- RQ2在稀疏比较图下,CARE 模型中 MLE 的最优 $\ell_2$ 与 $\ell_\infty$ 估计速率是什么?
- RQ3我们如何对 MLE 的不确定性进行量化,以涵盖基线得分 $\alpha_i^*$ 和协变量系数 $\bm{\beta}^*$?
- RQ4我们能否推导出 MLE 的渐近分布,以实现对协变量是否解释潜在得分的假设检验?
- RQ5CARE 模型中实现有效不确定性量化所需的最小样本复杂度是多少?
主要发现
- 在边概率为 $p$ 且每对项目有 $L$ 次比较的 Erdős-Rényi 比较图下,$\bm{\beta}^*$ 的 MLE 的 $\ell_2$ 估计误差界为 $\left\|\widehat{\bm{\beta}}_M - \bm{\beta}^*\right\|_2 \lesssim \kappa_1 \sqrt{\frac{(d+1)\log n}{npL}}$,且以高概率成立。
- $\alpha_i^*$ 的 MLE 在 $\ell_2$ 与 $\ell_\infty$ 范数下的速率分别为 $\widetilde{O}(\sqrt{\frac{1}{npL}})$,在稀疏图模型下达到极小最大风险的最优速率。
- 已建立 $\bm{\beta}^*$ 的 MLE 的渐近正态性,从而支持对单个协变量效应的置信区间构造与假设检验。
- 本文改进了 Gao et al. (2021) 对 BLT 模型的近似方法,并将其推广至 CARE 模型,显著提升了高维设置下的推断精度。
- 数值实验与对共同基金持股数据集的应用验证了理论结果,表明不确定性量化准确,且在对稀疏参数进行阈值化后排名性能得到提升。
- 带有 $\ell_2$ 正则化的投影梯度下降算法在 $\bm{\alpha}$ 上收敛可靠,实证结果证实了理论速率与推断质量。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。