[논문 리뷰] Learning Continuous User Representations through Hybrid Filtering with doc2vec
이 논문은 앱 사용 이력과 앱 메타데이터에서 연속적인 사용자 표현을 학습하기 위해 doc2vec를 사용하는 사용자2vec 및 컨텍스트2vec라는 새로운 하이브리드 필터링 방법을 제안한다. 사용자 행동과 콘텐츠 특징을 함께 모델링함으로써 룩알라인 모델링 성능이 크게 향상되며, 컨텍스트2vec는 하류 작업에서 직접적인 특징 공학보다 우수한 성능을 보이며 기준 모델 대비 상대 AUC-ROC 향상률 2.24%를 기록한다.
Players in the online ad ecosystem are struggling to acquire the user data required for precise targeting. Audience look-alike modeling has the potential to alleviate this issue, but models' performance strongly depends on quantity and quality of available data. In order to maximize the predictive performance of our look-alike modeling algorithms, we propose two novel hybrid filtering techniques that utilize the recent neural probabilistic language model algorithm doc2vec. We apply these methods to data from a large mobile ad exchange and additional app metadata acquired from the Apple App store and Google Play store. First, we model mobile app users through their app usage histories and app descriptions (user2vec). Second, we introduce context awareness to that model by incorporating additional user and app-related metadata in model training (context2vec). Our findings are threefold: (1) the quality of recommendations provided by user2vec is notably higher than current state-of-the-art techniques. (2) User representations generated through hybrid filtering using doc2vec prove to be highly valuable features in supervised machine learning models for look-alike modeling. This represents the first application of hybrid filtering user models using neural probabilistic language models, specifically doc2vec, in look-alike modeling. (3) Incorporating context metadata in the doc2vec model training process to introduce context awareness has positive effects on performance and is superior to directly including the data as features in the downstream supervised models.
연구 동기 및 목표
- 루크알라인 모델링을 위한 데이터 품질과 양을 향상시켜 온라인 광고에서 사용자 표현을 개선하기 위해.
- 실시간 입찰(RTB) 환경에서 제한된 사용자 데이터 문제를 앱 사용 이력과 메타데이터를 활용하여 해결하기 위해.
- neural probabilistic 언어 모델인 doc2vec가 예측 모델링을 위한 우수한 사용자 임베딩을 생성할 수 있는지 평가하기 위해.
- doc2vec 학습 과정에 맥락 메타데이터를 통합하는 것이 특징 연결보다 하류 성능 향상에 더 효과적인지 조사하기 위해.
제안 방법
- 사용자 앱 사용 시퀀스에 대해 doc2vec를 훈련시켜 연속적인 사용자 임베딩을 생성하는 하이브리드 필터링 방법인 user2vec를 제안한다.
- app 메타데이터(예: 장르, 가격, 평점 등)를 직접 doc2vec 훈련 과정에 통합함으로써 user2vec를 확장한 context2vec를 도입한다.
- 사용자-앱 상호작용 기반의 협업 필터링과 앱 설명 및 메타데이터 기반의 콘텐츠 기반 필터링을 하이브리드 모델링을 통해 통합한다.
- 비교를 위한 기준으로 TF-IDF를 사용하며, 이를 사용자 수준의 앱 설명에 적용하고 doc2vec 기반 표현과 결합한다.
- 성별 및 연령 분류 작업에서 예측 성능을 평가하기 위해 지도 학습 모델(예: XGBoost)을 활용한다.
- 밀도 높은 벡터 표현을 학습하기 위해 음성 샘플링과 스킵그램 아키텍처를 사용하여 doc2vec 모델을 훈련시킨다.
실험 결과
연구 질문
- RQ1R1: 앱 사용 이력과 설명에서 doc2vec 기반 user2vec 학습이 최신 기술 대비 추천 품질을 향상시킬 수 있는가?
- RQ2R2: doc2vec 훈련 과정에 앱 메타데이터를 통합하는 것(즉, context2vec)이 메타데이터를 별도의 특징으로 취급하는 것보다 더 나은 사용자 표현을 제공하는가?
- RQ3R3: 하이브리드 필터링 모델(예: TF-IDF + d2v:CF)의 성능은 성별 및 연령 예측을 위한 룩알라인 모델링에서 기준 모델 대비 어떻게 비교되는가?
주요 결과
- user2vec가 생성한 임베딩은 추천 품질에서 현재 최고 수준의 기술보다 뚜렷이 뛰어나며, 사용자 표현 학습 능력이 뛰어나다는 것을 입증한다.
- TF-IDF와 d2v:CF(협업 필터링)를 조합한 하이브리드 필터링 접근 방식은 기준 모델 대비 평균 상대 AUC-ROC 향상률 2.23%를 기록했다.
- 메타데이터를 doc2vec 훈련 과정에 통합하는 context2vec는 메타데이터를 직접 특징으로 포함시키는 것보다 하류 예측 성능 향상에 더 효과적이었으며, 기준 모델 대비 상대 AUC-ROC 향상률 2.24%를 기록했다.
- 가장 높은 성능을 보인 모델은 TF-IDF + d2v:CF + 메타데이터였으며, 이는 제로 기준 모델 대비 상대 AUC-ROC 향상률 2.24%와 절대 AUC-ROC 향상률 1.27%를 기록했다.
- doc2vec 훈련 과정에서 맥락 정보를 통합하는 것이 특징 공학보다 우수한 성능을 보였으며, 종단 간 맥락 인식 표현 학습의 이점이 입증되었다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.