Skip to main content
QUICK REVIEW

[論文レビュー] Uncertainty Quantification of MLE for Entity Ranking with Covariates

Jianqing Fan, Jikai Hou|arXiv (Cornell University)|Dec 20, 2022
Game Theory and Voting Systems被引用数 7
ひとこと要約

本稿では、アイテムの共変量を組み込んだラーニング・トゥ・ラーニング(CARE)モデルを提案する。これは、潜在スコア構造 $\alpha_i^* + \mathbf{x}_i^\top\bm{\beta}^*$ を通じて Bradley-Terry-Luce(BTL)モデルを拡張したものであり、最尤推定器(MLE)の $\ell_2$ および $\ell_\infty$ 評価レートを最適化し、不確実性の定量化のための漸近的分布を導出する。これにより、最小限のサンプル複雑性で共変量効果に関する統計的推論が可能になる。

ABSTRACT

This paper concerns with statistical estimation and inference for the ranking problems based on pairwise comparisons with additional covariate information such as the attributes of the compared items. Despite extensive studies, few prior literatures investigate this problem under the more realistic setting where covariate information exists. To tackle this issue, we propose a novel model, Covariate-Assisted Ranking Estimation (CARE) model, that extends the well-known Bradley-Terry-Luce (BTL) model, by incorporating the covariate information. Specifically, instead of assuming every compared item has a fixed latent score $\{θ_i^*\}_{i=1}^n$, we assume the underlying scores are given by $\{α_i^*+{x}_i^ opβ^*\}_{i=1}^n$, where $α_i^*$ and ${x}_i^ opβ^*$ represent latent baseline and covariate score of the $i$-th item, respectively. We impose natural identifiability conditions and derive the $\ell_{\infty}$- and $\ell_2$-optimal rates for the maximum likelihood estimator of $\{α_i^*\}_{i=1}^{n}$ and $β^*$ under a sparse comparison graph, using a novel `leave-one-out' technique (Chen et al., 2019) . To conduct statistical inferences, we further derive asymptotic distributions for the MLE of $\{α_i^*\}_{i=1}^n$ and $β^*$ with minimal sample complexity. This allows us to answer the question whether some covariates have any explanation power for latent scores and to threshold some sparse parameters to improve the ranking performance. We improve the approximation method used in (Gao et al., 2021) for the BLT model and generalize it to the CARE model. Moreover, we validate our theoretical results through large-scale numerical studies and an application to the mutual fund stock holding dataset.

研究の動機と目的

  • 共変量を組み込んだエンティティランク付けモデルに対する統計的推論手法の不足に応えること。これは現実世界で一般的な状況である。
  • ベースラインと共変量駆動スコアの和として潜在スコアをモデル化する、新たなモデル Covariate-Assisted Ranking Estimation(CARE)を構築すること。
  • スパースな比較グラフ下で、内在的スコア $\alpha_i^*$ と共変量係数 $\bm{\beta}^*$ の両方の MLE の最適な収束レートを確立すること。
  • MLE の漸近的正規性を導出し、共変量効果に関する信頼区間と仮説検定を可能にする。
  • 共変量が潜在スコアの変動を説明するかどうかを評価し、スパースなパラメータをしきい値処理してランク付け性能を向上させる、実用的な推論フレームワークを提供すること。

提案手法

  • CARE モデルを提案:$\mathbb{P}(\text{アイテム } i \text{ が } j \text{ を上回る}) = \frac{e^{\alpha_i^* + \mathbf{x}_i^\top\bm{\beta}^*}}{e^{\alpha_i^* + \mathbf{x}_i^\top\bm{\beta}^*} + e^{\alpha_j^* + \mathbf{x}_j^\top\bm{\beta}^*}}$。ここで $\alpha_i^*$ はベースライン、$\mathbf{x}_i^\top\bm{\beta}^*$ は共変量効果を表す。
  • 一意な推定を保証するため、識別可能性制約を課した制約付き最尤推定器(MLE)$\widehat{\bm{\theta}}_M = (\widehat{\bm{\alpha}}_M, \widehat{\bm{\beta}}_M)$ を用いる。
  • スパースな Erdős-Rényi 比較グラフ下で、$\widehat{\bm{\alpha}}_M$ および $\widehat{\bm{\beta}}_M$ の $\ell_2$ および $\ell_\infty$ 収束レートを導出するために、新規の「リーブ・オブ・ワン・アウト」技術(Chen et al., 2019)を適用する。
  • MLE の $\bm{\beta}^*$ および $\alpha_i^*$ の漸近的正規性を導出し、個々の共変量効果に関する信頼区間と仮説検定を可能にする。
  • Gao et al. (2021) における BLT モデルの近似手法を改善し、それらを CARE モデルに一般化することで、高次元設定における推論精度を向上させる。
  • 数値最適化の収束性と安定性を確保するため、$\bm{\alpha}$ に $\ell_2$ 正則化を施した投影勾配降下法を実装する。

実験結果

リサーチクエスチョン

  • RQ1スパースな比較グラフ下で、推定の有効性が保証され、アイテムの共変量を組み込んだペアワイズ比較ランク付けに統計的推論が可能なモデルを設計できるか?
  • RQ2スパースな比較グラフ下で、CARE モデルにおける MLE の最適な $\ell_2$ および $\ell_\infty$ 評価レートは何か?
  • RQ3MLE の不確実性を、潜在スコア $\alpha_i^*$ および共変量係数 $\bm{\beta}^*$ の両方についてどのように定量化できるか?
  • RQ4共変量が潜在スコアを説明するかどうかを検証するための仮説検定を可能にする、MLE の漸近的分布を導出できるか?
  • RQ5CARE モデルで有効な不確実性の定量化を達成するために必要な最小サンプル複雑性は何か?

主な発見

  • Erdős-Rényi 比較グラフ(エッジ確率 $p$、各ペアあたり $L$ 回の比較)下で、$\bm{\beta}^*$ の MLE は高確率で $\left\|\widehat{\bm{\beta}}_M - \bm{\beta}^*\right\|_2 \lesssim \kappa_1 \sqrt{\frac{(d+1)\log n}{npL}}$ の $\ell_2$ 評価誤差バウンドを達成する。
  • $\alpha_i^*$ の MLE は $\ell_2$ および $\ell_\infty$ レートとして $\widetilde{O}(\sqrt{\frac{1}{npL}})$ を達成し、スパースグラフモデル下でのミニマックス最適レートと一致する。
  • $\bm{\beta}^*$ の MLE の漸近的正規性が確立され、個々の共変量効果に関する信頼区間と仮説検定の構築が可能になる。
  • 本稿は、Gao et al. (2021) における BLT モデルの近似手法を改善し、それを CARE モデルに一般化することで、高次元設定における推論精度を向上させた。
  • 数値実験およびマーケット・ファンドの株式保有データセットへの応用により、理論的結果が妥当であることが検証され、不確実性の定量化が正確であり、スパースパラメータのしきい値処理によるランク付け性能の向上が確認された。
  • 投影勾配降下法に $\ell_2$ 正則化を施したアルゴリズムは信頼性高く収束し、実験的結果が理論的レートおよび推論品質を裏付けている。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。