[논문 리뷰] Intent term selection and refinement in e-commerce queries
이 논문은 월마트의 쿼리 재구성 로그를 활용하여 전자상거래 검색을 위한 맥락 인식 텀 가중치 및 쿼리 정련 방법을 제안한다. RNN 기반의 의도 인코딩과 맥락 기반 텀 표현을 활용함으로써 희귀 쿼리의 관련성을 향상시키며, 의도 핵심 텀을 식별하고 어휘 갭을 메우는 텀을 제안함으로써, MRR 및 정밀도 메트릭에서 비맥락 기반 베이스라인을 능가한다.
In e-commerce, a user tends to search for the desired product by issuing a query to the search engine and examining the retrieved results. If the search engine was successful in correctly understanding the user's query, it will return results that correspond to the products whose attributes match the terms in the query that are representative of the query's product intent. However, the search engine may fail to retrieve results that satisfy the query's product intent and thus degrading user experience due to different issues in query processing: (i) when multiple terms are present in a query it may fail to determine the relevant terms that are representative of the query's product intent, and (ii) it may suffer from vocabulary gap between the terms in the query and the product's description, i.e., terms used in the query are semantically similar but different from the terms in the product description. Hence, identifying the terms that describe the query's product intent and predicting additional terms that describe the query's product intent better than the existing query terms to the search engine is an essential task in e-commerce search. In this paper, we leverage the historical query reformulation logs of a major e-commerce retailer to develop distant-supervised approaches to solve both these problems. Our approaches exploit the fact that the significance of a term is dependent upon the context (other terms in the neighborhood) in which it is used in order to learn the importance of the term towards the query's product intent. We show that identifying and emphasizing the terms that define the query's product intent leads to a 3% improvement in ranking. Moreover, for the tasks of identifying the important terms in a query and for predicting the additional terms that represent product intent, experiments illustrate that our approaches outperform the non-contextual baselines.
연구 동기 및 목표
- 기본적으로는 역사적 참여 데이터가 제한된 희귀 또는 콜드스타트 쿼리에 대해 전자상거래 검색의 관련성을 향상시키는 것.
- 특히 혼란스럽거나 노이즈가 많은 단어가 포함된 경우, 쿼리 내에서 사용자의 제품 의도를 가장 정확하게 표현하는 단어를 식별하는 것.
- 사용자 쿼리와 제품 카탈로그 단어 간의 어휘 갭을 메우기 위해 더 관련성 있고 맥락적으로 적절한 단어를 제안하는 것.
- 명시적 애너테이션에 의존하지 않고 쿼리 재구성 패턴을 활용해 의도를 추론할 수 있는 일반화 가능한 방법을 개발하는 것.
제안 방법
- 월마트의 역사적 쿼리 재구성 로그를 활용해 의도 탐지용 원거리 지도 학습 모델을 훈련한다.
- 쿼리 텀의 맥락적 표현을 모델링하기 위해 순환 신경망(RNN) 인코더를 사용한다. 이는 단어의 중요성이 주변 엔티티에 따라 어떻게 변화하는지 파악한다.
- 맥락 기반 임베딩에 기반해 쿼리의 제품 의도를 가장 잘 표현하는 텀에 더 높은 가중치를 할당하는 텀 가중치 모델(CTW)을 도입한다.
- 의도 인코더와 다중 레이블 분류기를 사용해 원래 쿼리에 포함되지 않은 관련 텀을 예측하는 쿼리 정련 모델(CQR)을 개발한다.
- 텀 중요도 추정과 어휘 갭 해소를 결합하여 제품 카탈로그 언어와 더 잘 맞는 텀을 예측한다.
- 희귀 쿼리의 검증 세트에서 MRR 및 정밀도 메트릭을 사용해 모델을 평가하며, TF-IDF, FTW, VPCG, VG 등의 베이스라인과 비교한다.
실험 결과
연구 질문
- RQ1의도가 모호하거나 노이즈가 많은 단어가 포함된 쿼리에서 사용자의 진짜 제품 의도를 표현하는 데 가장 관련성이 높은 단어를 어떻게 식별할 수 있는가?
- RQ2이웃하는 단어의 맥락 정보를 통합함으로써 비맥락 기반 방법에 비해 텀 가중치 정확도가 얼마나 향상되는가?
- RQ3쿼리 재구성 패턴을 활용해 사용자 쿼리와 제품 카탈로그 설명어 사이의 어휘 갭을 메우는 새로운 효과적인 단어를 제안할 수 있는가?
- RQ4희귀 쿼리에 대해 역사적 참여 데이터가 부족한 상황에서 맥락 인식 모델은 얼마나 효과적으로 검색 성능을 향상시키는가?
주요 결과
- 맥락 인식 텀 가중치 모델(CTW)은 MRR 및 정밀도 측정 기준으로 비맥락 기반 베이스라인(TF-IDF, FTW, VPCG, VG)보다 높은 순위 관련성을 보이며 성능이 뛰어나다.
- 쿼리 'battery night light with timer'의 경우 CTW는 'night'과 'light'에 가장 높은 가중치를 할당하는 데 성공했으며, 베이스라인은 잘못되게 'timer'나 'battery'를 우선시한다.
- 쿼리 정련 모델(CQR)은 'orbit red garden hose water nozzle'에 대해 'water spray nozzle'과 같이 맥락적으로 관련성 있는 텀을 성공적으로 예측했으며, 'gum'이나 'spearmint'와 같은 부적절한 연관성을 피했다.
- 'auto seat cover wonder woman'의 경우 CTW는 'auto', 'seat', 'cover'를 핵심 텀으로 정확히 식별했고, 베이스라인은 의도를 해석하지 못해 오류를 범했다.
- CQR 모델은 의도 인코더를 통해 제품 유형을 이해함으로써 비맥락 기반의 FQR 베이스라인과 달리 관련 없는 단어를 생성하지 않아, 맥락을 고려한 정교한 추론이 가능하다.
- 기존 방법이 역사적 참여 데이터 부족으로 실패하는 데 비해, 이 방법은 희귀 쿼리에서 뛰어난 성능을 보이며 콜드스타트 시나리오에서의 가치를 입증한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.