[論文レビュー] A Nonparametric Delayed Feedback Model for Conversion Rate Prediction
この論文は、指数分布やワイブル分布のようなパラメトリックな形を仮定せずに、時間遅延分布を推定する非パラメトリックな遅延フィードバックモデルであるNoDeFを提案する。カーネル密度推定と学習された重み、EMアルゴリズムを用いることで、複雑でデータ駆動の遅延分布を捉え、実世界のCriteoデータにおいて、ログ損失、正解率、AUCの観点で従来のパラメトリックモデルを上回る性能を発揮する。
Predicting conversion rates (CVRs) in display advertising (e.g., predicting the proportion of users who purchase an item (i.e., a conversion) after its corresponding ad is clicked) is important when measuring the effects of ads shown to users and to understanding the interests of the users. There is generally a time delay (i.e., so-called {\it delayed feedback}) between the ad click and conversion. Owing to the delayed feedback, samples that are converted after an observation period may be treated as negative. To overcome this drawback, CVR prediction assuming that the time delay follows an exponential distribution has been proposed. In practice, however, there is no guarantee that the delay is generated from the exponential distribution, and the best distribution with which to represent the delay depends on the data. In this paper, we propose a nonparametric delayed feedback model for CVR prediction that represents the distribution of the time delay without assuming a parametric distribution, such as an exponential or Weibull distribution. Because the distribution of the time delay is modeled depending on the content of an ad and the features of a user, various shapes of the distribution can be represented potentially. In experiments, we show that the proposed model can capture the distribution for the time delay on a synthetic dataset, even when the distribution is complicated. Moreover, on a real dataset, we show that the proposed model outperforms the existing method that assumes an exponential distribution for the time delay in terms of conversion rate prediction.
研究の動機と目的
- 広告クリックとコンバージョンの間の時間遅延に、指数分布などのパラメトリックな分布を仮定する従来のCVR予測モデルの限界を是正すること。
- 複雑でデータ依存の形状を柔軟に表現できる非パラメトリックな方法で、時間遅延分布をモデル化すること。
- 遅延フィードバックを正確に捉えることで、遅延分布に関する事前の仮定に依存せずに予測性能を向上させること。
- 表示広告に限らず、遅延フィードバックが生じるあらゆる状況に適用可能な汎用的なフレームワークを開発すること。
提案手法
- NoDeFは、観測されたコンバージョン時刻および時間軸上の擬似点を中心に配置されたカーネル関数の重み付き和を用いて、時間遅延分布をモデル化する。
- カーネルの重みは、隠れたコンバージョン状態とモデルパラメータを同時に推定するEMアルゴリズムを用いてデータから学習する。
- この手法はカーネル密度推定の原則を採用しているが、遅延フィードバック設定における打ち切りデータと隠れ変数を扱えるように拡張している。
- ハイパーパrameterとして、擬似点の数(L)、正則化項(λw、λV)は検証データ上で最適化される。
- カテゴリカル特徴量のワンホットエンコーディングと遅延時間の対数変換後の正規化の後、PCAを用いて特徴量表現を100次元に削減する。
- 大規模なコンバージョンログへのスケーラビリティを実現するため、確率的EMを用いてモデルを学習する。
実験結果
リサーチクエスチョン
- RQ1特定のパラメトリックな形を仮定せずに、コンバージョンイベントにおける時間遅延の真の分布を非パラメトリックに効果的に推定できるか?
- RQ2パラメトリックモデルと比較して、NoDeFは複雑で多峰的、または非指数的な遅延分布をどれほど正確に捉えられるか?
- RQ3非パラメトリックアプローチにより、実世界のデータセットにおけるCVR推定の予測性能が向上するか?
- RQ4モデルは未観測のサンプルに一般化可能であり、未コンバージョン(打ち切り)データに対しても効果的に対処できるか?
主な発見
- 合成データセットでは、真の分布が非指数的で複雑であっても、NoDeFは二峰性の時間遅延分布(二つの明確なピーク)を的確に捉えた。
- Criteoデータセットでは、ログ損失、正解率、AUCのすべての評価指標において、NAIVEベースラインとDFMモデルを上回った。
- 最近のキャンペーンサブセットではAUCが0.781、全キャンペーンサブセットでは0.763を達成し、DFMとNAIVEを上回った。
- L(10、20)の異なる値における推定密度曲線は、滑らかでデータに適応した形状を示し、生データの二峰性構造と一致しており、非パラメトリック推定の有効性を示している。
- NoDeFにおけるカーネル密度推定のための自動帯域幅選択法は、さまざまな設定において安定的かつ信頼性の高い密度推定を実現した。
- 結果から、真の遅延分布が複雑または標準的でない場合、指数分布などのパラメトリックな遅延分布を仮定することはモデル性能を制限することが示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。