Skip to main content
QUICK REVIEW

[論文レビュー] Validation Set Evaluation can be Wrong: An Evaluator-Generator Approach for Maximizing Online Performance of Ranking in E-commerce

Guangda Huzhang, Zhen-Jia Pang|arXiv (Cornell University)|Mar 25, 2020
Mobile Crowdsensing and Crowdsourcing参考文献 30被引用数 5
ひとこと要約

本論文は、eコマースの学習ランキングシステムにおけるオフライン評価の信頼性を向上させる、評価者・生成者フレームワークを提案する。オフラインデータ上で評価者を訓練し、オンライン性能を模倣するようにし、強化学習を用いて生成者をその評価者に従って最適化することで、実際のオンライン性能をよりよく予測する検証スコアを達成した。その結果、AliExpress Searchにおいて、コンversion rateとGMVが2%以上向上した。

ABSTRACT

Learning-to-rank (LTR) has become a key technology in E-commerce applications. Most existing LTR approaches follow a supervised learning paradigm, relying on labeled data collected from a specific online system. However, the online performance of the models may be inconsistent with that observed in offline evaluation. This inconsistency is serious: even if we successfully reproduce the offline performance of a newly proposed model, it may perform poorly when deployed in online systems. Interacting with online systems is an accurate albeit costly way of evaluation as the process may hurt user experience. To avoid the inaccuracy of offline evaluation and the cost of the online interaction-based evaluation, we propose an evaluator-generator framework. Firstly, we train an evaluator model as an approximation of online systems with offline data. Secondly, we learn a generator with supervision from the evaluator, which can approximate the maximal online performance through reinforcement learning. Through extensive experience, we show that the classic data-based metrics on the validation set can be inconsistent with online performance, and can even be misleading. We also demonstrate that the proposed evaluator score is significantly more robust than common ranking metrics: classic metrics do not match the actual performance in both an offline simulated environment and a real online system while our evaluator score matches them well. Finally, we show that our method achieves a significant improvement of ($ extgreater2\%$) over the current industrial-level pair-wise model in terms of both Conversion Rate (CR) and Gross Merchandise Volume (GMV) in online A/B tests on AliExpress Search. Roughly, it contributes more than 300 millions dollars to GMV per year in the common daily selling.

研究の動機と目的

  • eコマースの学習ランキングシステムにおけるオフライン検証パフォーマンスと実際のオンライン順序付けパフォーマンスの不一致を解消すること。
  • オフライン評価の正確性を向上させることで、高コストでリスクの高いオンラインA/Bテストへの依存を減らすこと。
  • オフラインデータと強化学習のみを用いて、最大のオンラインパフォーマンスを予測する手法を開発すること。
  • 従来のオフライン指標が順序付けモデルの評価において誤解を招く可能性があることを示すこと。
  • シミュレートされたオフライン環境とAliExpress Searchでの実際のオンラインA/Bテストの両方を用いて、提案手法の有効性を検証すること。

提案手法

  • オフラインラベル付きデータを用いて評価者モデルを訓練し、実際のオンラインシステムの挙動を近似する。
  • 訓練済みの評価者を強化学習フレームワーク内の報酬信号として用い、生成者モデルを訓練する。
  • 生成者は、評価者が予測するオンラインパフォーマンスを最大化するように最適化される。
  • 評価者スコアはオンラインパフォーマンスの代理指標として機能し、信頼性の高いオフラインモデル選択を可能にする。
  • オフラインおよびオンライン評価の両方で比較の基準として、ペアワイズ学習ランキングモデルをベースラインとして用いる。
  • 生成者を用いて最終的な順序付けモデルを生成し、実際のAliExpress Searchシステム上でA/Bテストにより評価する。

実験結果

リサーチクエスチョン

  • RQ1従来のオフライン順序付け指標(検証セット上)は、eコマース順序付けシステムにおけるオンラインパフォーマンスを信頼性を持って予測できるか?
  • RQ2提案された評価者・生成者フレームワークは、オフライン評価と実際のオンラインパフォーマンスの整合性をどの程度向上させるか?
  • RQ3評価者スコアは、NDCG や MAP といった標準指標と比較して、現実のオンライン結果をどの程度よく予測できるか?
  • RQ4提案手法は、実際のオンライン展開において、コンversion rate や GMV といった重要なビジネス指標を顕著に改善できるか?
  • RQ5評価者スコアは、シミュレートされたオフライン環境と実際のオンラインシステムの両方で頑健か?

主な発見

  • 検証セット上の古典的オフライン指標(NDCG や MAP)は、実際のオンラインパフォーマンスと整合性がなく、誤解を招く可能性があることが判明した。
  • 提案された評価者スコアは、シミュレートされたオフライン環境および実際のオンラインシステムの両方の結果と強く整合しており、標準指標よりも顕著に高い頑健性を示した。
  • 実際のオンラインA/Bテストにおいて、コンversion rate およびグロスマーチャンダイズボリューム(GMV)の両方で、統計的に有意な2%以上の改善が達成された。
  • 日常の売上高を基に換算すると、提案手法によるGMVへの年間貢献額は3億ドル以上にのぼると推定された。
  • 評価者・生成者フレームワークは、オフラインデータと強化学習のみを用いて、最大の達成可能なオンラインパフォーマンスを成功裏に近似できた。
  • 結果から、従来の指標を用いたオフライン評価は、実際のオンライン成功を予測できない可能性がある一方で、提案された評価者スコアは信頼できる代替手段であることが確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。