Skip to main content
QUICK REVIEW

[論文レビュー] Intent term selection and refinement in e-commerce queries

Saurav Manchanda, Mohit Sharma|arXiv (Cornell University)|Aug 22, 2019
Web Data Mining and Analysis参考文献 21被引用数 4
ひとこと要約

本稿では、ウォーマートのクエリ再定式化ログを活用して、eコマース検索における文脈に配慮した用語重み付けおよびクエリ精錬手法を提案する。RNNベースの意図符号化と文脈的用語表現を活用することで、希少なクエリの検索関連性が向上し、意図に重要な用語を特定するとともに、語彙ギャップを埋める用語を提案する。MRRおよび精度指標において、非文脈的ベースラインを上回る性能を発揮する。

ABSTRACT

In e-commerce, a user tends to search for the desired product by issuing a query to the search engine and examining the retrieved results. If the search engine was successful in correctly understanding the user's query, it will return results that correspond to the products whose attributes match the terms in the query that are representative of the query's product intent. However, the search engine may fail to retrieve results that satisfy the query's product intent and thus degrading user experience due to different issues in query processing: (i) when multiple terms are present in a query it may fail to determine the relevant terms that are representative of the query's product intent, and (ii) it may suffer from vocabulary gap between the terms in the query and the product's description, i.e., terms used in the query are semantically similar but different from the terms in the product description. Hence, identifying the terms that describe the query's product intent and predicting additional terms that describe the query's product intent better than the existing query terms to the search engine is an essential task in e-commerce search. In this paper, we leverage the historical query reformulation logs of a major e-commerce retailer to develop distant-supervised approaches to solve both these problems. Our approaches exploit the fact that the significance of a term is dependent upon the context (other terms in the neighborhood) in which it is used in order to learn the importance of the term towards the query's product intent. We show that identifying and emphasizing the terms that define the query's product intent leads to a 3% improvement in ranking. Moreover, for the tasks of identifying the important terms in a query and for predicting the additional terms that represent product intent, experiments illustrate that our approaches outperform the non-contextual baselines.

研究の動機と目的

  • 歴史的エンゲージメントデータが限られる希少またはコールドスタートクエリにおけるeコマース検索の関連性を向上させること。
  • ノイズが多いか曖昧な用語を含むクエリにおいても、ユーザーの商品意図を最も正確に表現する用語を特定すること。
  • ユーザークエリと商品カタログ用語の間の語彙ギャップを埋めるために、より関連性が高く文脈的に適切な用語を提案すること。
  • 明示的なアノテーションに依存せずに、クエリ再定式化パターンを活用して意図を推定する汎用的な手法を開発すること。

提案手法

  • ウォーマートの歴史的クエリ再定式化ログを活用し、意図検出のための遠隔教師付きモデルを学習する。
  • 再帰的ニューラルネットワーク(RNN)エンコーダーを用いて、用語の文脈的表現をモデル化し、周囲のエンティティに応じた用語の重要性の変化を捉える。
  • 文脈に配慮した埋め込みに基づき、クエリの商品意図を最もよく表現する用語に高い重みを付与する、用語重み付けモデル(CTW)を導入する。
  • 意図エンコーダーとマルチラベル分類器を用いて、元のクエリに存在しない関連する用語を予測するクエリ精錬モデル(CQR)を開発する。
  • 用語の重要度推定と語彙ギャップ解消を統合し、商品カタログの言語により適切に一致する用語を予測する。
  • 希少クエリのホールドアウトセットを用いてMRRおよび精度指標でモデルを評価し、TF-IDF、FTW、VPCG、VGなどのベースラインと比較する。

実験結果

リサーチクエスチョン

  • RQ1曖昧またはノイズの多い用語を含むクエリにおいて、真の商品意図を最もよく表現する用語をどのように特定できるか?
  • RQ2周囲の用語からの文脈的情報を組み込むことで、非文脈的メソッドと比較して用語重み付けの正確性がどの程度向上するか?
  • RQ3クエリ再定式化パターンを活用して、ユーザークエリと商品カタログ記述語の間の語彙ギャップを埋める、より効果的な新しい用語を提案できるか?
  • RQ4文脈に配慮したモデルは、歴史的エンゲージメントデータが限られる希少クエリの検索パフォーマンス向上にどの程度効果的か?

主な発見

  • 文脈に配慮した用語重み付けモデル(CTW)は、TF-IDF、FTW、VPCG、VGなどの非文脈的ベースラインを上回り、MRRおよび精度指標で高い関連性ランキングを達成した。
  • クエリ「battery night light with timer」において、CTWは「night」と「light」に最も高い重みを付与したが、ベースラインは「timer」や「battery」を誤って優先した。
  • クエリ精錬モデル(CQR)は、「orbit red garden hose water nozzle」に対して「water spray nozzle」などの文脈的に適切な用語を正しく予測したが、誤った関連付けである「gum」や「spearmint」といった用語は避けていた。
  • 「auto seat cover wonder woman」のケースでは、CTWは「auto」、「seat」、および「cover」をキーワードとして正しく特定したが、ベースラインは意図を明確に特定できなかった。
  • CQRモデルは、製品タイプを理解するための意図エンコーダーを活用することで、不要な用語の生成を回避した。これに対してFQRベースラインは文脈を欠いているため、誤った提案を生じていた。
  • 本手法は希少クエリにおいて強く性能を発揮し、従来の手法が歴史的エンゲージメントデータの不足により失敗する状況でも有効であることが確認され、コールドスタートシナリオにおける価値を裏付けた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。