Skip to main content
QUICK REVIEW

[논문 리뷰] User Response Prediction in Online Advertising

Zhabiz Gharibshah, Xingquan Zhu|arXiv (Cornell University)|2021. 01. 07.
Recommender Systems and Techniques참고 문헌 192인용 수 9
한 줄 요약

이 종합적 서베이는 온라인 광고에서 사용자 반응 예측을 위한 기계학습 및 딥러닝 방법에 대한 포괄적인 분류 체계를 제시한다. 플랫폼, 이해관계자, 데이터 유형, 기술적 접근 방식을 포함한다. 최신 모델, 벤치마크 데이터셋, 오픈소스 구현을 검토하여 클릭-through rate 예측 및 사용자 행동 모델링 분야의 산업 및 학술 연구를 안내한다.

ABSTRACT

Online advertising, as the vast market, has gained significant attention in various platforms ranging from search engines, third-party websites, social media, and mobile apps. The prosperity of online campaigns is a challenge in online marketing and is usually evaluated by user response through different metrics, such as clicks on advertisement (ad) creatives, subscriptions to products, purchases of items, or explicit user feedback through online surveys. Recent years have witnessed a significant increase in the number of studies using computational approaches, including machine learning methods, for user response prediction. However, existing literature mainly focuses on algorithmic-driven designs to solve specific challenges, and no comprehensive review exists to answer many important questions. What are the parties involved in the online digital advertising eco-systems? What type of data are available for user response prediction? How to predict user response in a reliable and/or transparent way? In this survey, we provide a comprehensive review of user response prediction in online advertising and related recommender applications. Our essential goal is to provide a thorough understanding of online advertising platforms, stakeholders, data availability, and typical ways of user response prediction. We propose a taxonomy to categorize state-of-the-art user response prediction methods, primarily focus on the current progress of machine learning methods used in different online platforms. In addition, we also review applications of user response prediction, benchmark datasets, and open-source codes in the field.

연구 동기 및 목표

  • 온라인 광고 생태계를 체계적으로 이해하기 위해 광고주, 게재자, 플랫폼과 같은 핵심 이해관계자를 포함한 생태계의 구조를 제시한다.
  • 다양한 광고 플랫폼에서 사용자 반응 예측에 사용되는 데이터 및 특징의 유형을 식별하고 분류한다.
  • 클릭, 전환, 구매와 같은 사용자 반응을 예측하기 위한 최신 기계학습 및 딥러닝 방법의 체계적 분류 체계를 수립한다.
  • 실제 산업 응용 사례, 벤치마크 데이터셋, 오픈소스 구현을 검토하여 재현 가능성과 향후 연구를 지원한다.
  • 실제 구현 환경에서 사용자 반응 예측 시스템의 신뢰성, 투명성, 확장성에 대한 주요 과제를 다룬다.

제안 방법

  • 모델 아키텍처 기반으로 사용자 반응 예측 방법을 분류하기 위한 분류 체계를 제안한다. 인과력 기반 모델, 딥 네트워크, 그래프 신경망, 하이브리드 모델을 포함한다.
  • DeepFM, DIN, DIEN, PIN, DCN_V2와 같은 대표적 모델을 검토하며, 사용자-아이템 상호작용 및 사용자 관심의 진화를 학습하는 메커니즘을 강조한다.
  • Facebook의 FBCTR, Etsy의 EtsyCTR, Google의 DCN_V2와 같은 산업 규모의 시스템을 분석하며, 특징 해싱, 모델 병렬화, 분산 학습과 같은 확장성 확보 기법을 중심으로 다룬다.
  • PinSage와 RippleNet와 같은 그래프 기반 모델을 검토하며, 사용자-아이템 상호작용 그래프를 활용해 표현 학습 성능을 향상시킨다.
  • 협업 필터링, 콘텐츠 기반 특징, 어텐션 메커니즘의 조합을 통해 동적 사용자 선호도를 모델링하는 하이브리드 접근 방식을 소개한다.
  • 기억 네트워크(MIMN, SIM)와 자기 어텐션 모듈을 활용한 기법을 평가하여 장기적인 사용자 행동 시퀀스 처리 및 관련성 점수 산정을 수행한다.
Figure 1. Advertising Eco-system. From left to right, a process is triggered when users start to interact with online services through either visiting a web-page, searching an item, or checking the social media in publisher website. In the case that the web-page has web placement available, the publ
Figure 1. Advertising Eco-system. From left to right, a process is triggered when users start to interact with online services through either visiting a web-page, searching an item, or checking the social media in publisher website. In the case that the web-page has web placement available, the publ

실험 결과

연구 질문

  • RQ1온라인 광고 생태계의 핵심 구성요소와 이해관계자는 무엇이며, 사용자 반응 예측 과정에서 어떻게 상호작용하는가?
  • RQ2검색, 디스플레이, 인앱 광고 등 다양한 광고 플랫폼에서 사용자 반응 모델링을 위해 사용 가능한 데이터 및 특징의 유형은 무엇인가?
  • RQ3현대 기계학습 및 딥러닝 모델은 클릭-through rate 및 전환 예측의 정확도와 효율성을 어떻게 향상시키는가?
  • RQ4대규모로 사용자 반응 예측 모델을 구현할 때 발생하는 주요 기술적 과제는 무엇이며, 산업 시스템에서는 어떻게 해결하는가?
  • RQ5기존 방법에 비해 하이브리드 및 그래프 기반 모델은 사용자 관심 모델링과 장기적 행동 이해도를 어떻게 향상시키는가?

주요 결과

  • 이 서베이는 사용자 반응 예측이 대부분 클릭-through rate(CTR) 추정 문제로 설정되며, 딥러닝 모델이 전통적 방법에 비해 정확도와 적응성 면에서 뛰어나다는 점을 규명한다.
  • PinSage 및 RippleNet와 같은 그래프 신경망(GNN)은 복잡한 사용자-아이템 상호작용 그래프를 모델링함으로써 추천 및 디스플레이 광고 분야에서 성능 향상을 이룬다.
  • Facebook의 FBCTR 및 Google의 DCN_V2와 같은 산업용 시스템은 서브샘플링, 모델 병렬화, 파라미터 공유 기법을 통해 고규모 확장성을 확보하여 실시간 추론을 가능하게 한다.
  • DIEN 및 SIM과 같은 모델은 어텐션 메커니즘과 기억 네트워크를 통합하여 변화하는 사용자 관심과 장기적인 순차적 행동을 포착함으로써 동적 사용자 프로파일에 대한 예측 성능을 향상시킨다.
  • DeepFM 및 PIN과 같은 하이브리드 모델은 인과력 기반 모델과 딥 네트워크를 조합하여 저차수 및 고차수 특징 상호작용을 효과적으로 학습하며, 벤치마크 데이터셋에서 최신 성능을 달성한다.
  • 이 서베이는 Criteo, Avazu, MovieLens와 같은 벤치마크 데이터셋과 오픈소스 구현체가 재현 가능성 및 분야 발전을 위한 핵심 요소임을 강조한다. 또한 PyTorch 및 Caffe2와 같은 도구는 모델 배포에 널리 사용된다.
Figure 2. The schema of user response prediction workflow. Embedding layer is the common paradigm to deal with high dimensional binary representation in user response prediction. They can either be set by pre-defined values or be trained as internal parameters in end-to-end models like deep learning
Figure 2. The schema of user response prediction workflow. Embedding layer is the common paradigm to deal with high dimensional binary representation in user response prediction. They can either be set by pre-defined values or be trained as internal parameters in end-to-end models like deep learning

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.