Skip to main content
QUICK REVIEW

[논문 리뷰] Evaluating Interpolation and Extrapolation Performance of Neural Retrieval Models

Jingtao Zhan, Xiaohui Xie|arXiv (Cornell University)|2022. 04. 25.
Domain Adaptation and Few-Shot Learning인용 수 4
한 줄 요약

이 논문은 쿼리 임bedding 유사도 기반으로 훈련 및 테스트 데이터를 재표본화하여 신경 검색 모델의 보간 및 외삽 성능을 별도로 평가할 수 있는 새로운 평가 프로토콜을 제안한다. 결과적으로 표현 기반 모델은 보간에서는 상호작용 기반 모델과 거의 유사한 성능를 보이지만, 외삽에서는 유의미하게 열등한 성능를 보이며, 표준 벤치마크를 초월한 일반화 능력에 대한 별도의 평가가 필요함을 시사한다.

ABSTRACT

A retrieval model should not only interpolate the training data but also extrapolate well to the queries that are different from the training data. While neural retrieval models have demonstrated impressive performance on ad-hoc search benchmarks, we still know little about how they perform in terms of interpolation and extrapolation. In this paper, we demonstrate the importance of separately evaluating the two capabilities of neural retrieval models. Firstly, we examine existing ad-hoc search benchmarks from the two perspectives. We investigate the distribution of training and test data and find a considerable overlap in query entities, query intent, and relevance labels. This finding implies that the evaluation on these test sets is biased toward interpolation and cannot accurately reflect the extrapolation capacity. Secondly, we propose a novel evaluation protocol to separately evaluate the interpolation and extrapolation performance on existing benchmark datasets. It resamples the training and test data based on query similarity and utilizes the resampled dataset for training and evaluation. Finally, we leverage the proposed evaluation protocol to comprehensively revisit a number of widely-adopted neural retrieval models. Results show models perform differently when moving from interpolation to extrapolation. For example, representation-based retrieval models perform almost as well as interaction-based retrieval models in terms of interpolation but not extrapolation. Therefore, it is necessary to separately evaluate both interpolation and extrapolation performance and the proposed resampling method serves as a simple yet effective evaluation tool for future IR studies.

연구 동기 및 목표

  • 기존의 임시적 검색 벤치마크가 진정으로 일반화 능력을 평가하는지 아니면 보간에 치우쳐 있는지 조사하기 위해.
  • 새로운 주석이나 테스트 세트가 필요 없이 보간 및 외삽 성능를 분리하고 측정할 수 있는 방법을 개발하기 위해.
  • 보간 및 외삽 환경에서 별도로 작동하는 다양한 신경 검색 모델 아키텍처의 일반화 행동을 이해하기 위해 널리 사용되는 신경 검색 모델을 재평가하기 위해.
  • 외삽 성능가 OOD(Out-of-Distribution) 일반화와 강하게 상관되며, 반대로 보간 성능는 그렇지 않다는 것을 입증하기 위해.

제안 방법

  • 기존 벤치마크(예: MS MARCO, TREC DL)의 쿼리를 사전 학습된 모델을 사용해 벡터 공간에 임bedding한다.
  • 임베딩 코사인 거리 기반 유사도를 계산하여, 유사한 쿼리(보간) 또는 비유사한 쿼리(외삽)로 훈련-테스트 쌍을 정의한다.
  • 데이터셋을 재표본화하여 새로운 훈련 및 테스트 분할을 생성한다: 하나는 유사한 쿼리를 강조하는 보간, 다른 하나는 비유사한 쿼리를 강조하는 외삽.
  • 모델을 이러한 재표본화된 분할에서 훈련하고 평가하여 각 환경에서의 성능를 고립시킨다.
  • 외삽 성능가 별도의 테스트 세트에서 OOD 일반화 성능와 비교하여 검증하였으며, 강한 상관관계를 보였다.

실험 결과

연구 질문

  • RQ1현재의 임시적 검색 벤치마크가 진정으로 외삽을 평가하는지 아니면 보간에 치우쳐 있는가?
  • RQ2다양한 신경 검색 모델 아키텍처(예: 표현 기반 vs. 상호작용 기반)가 보간 및 외삽 환경에서 어떻게 성능를 내는가?
  • RQ3제안된 재표본화 기반 평가 프로토콜이 외삽 성능를 효과적으로 고립하고 측정하는가?
  • RQ4재표본화된 데이터에서의 외삽 성능가 보존된 테스트 세트에서의 OOD 일반화 성능와 얼마나 잘 상관되는가?

주요 결과

  • MS MARCO 및 TREC Deep Learning Tracks와 같은 기존 벤치마크는 훈련 및 테스트 세트 간에 쿼리 엔티티, 의도, 관련성 레이블에 있어 상당한 중복을 보이며, 이는 보간에 치우친 강한 편향을 시사한다.
  • 표현 기반 모델(예: 디ensa retrieval)은 보간에서는 정점 성능를 기록하지만, ColBERT와 같은 상호작용 기반 모델에 비해 외삽에서 성능가 유의미하게 떨어진다.
  • 재표본화된 데이터에서의 외삽 성능가 OOD 일반화 성능와 강하게 상관되며, 이는 방법의 효과성을 검증한다.
  • 보간 성능는 OOD 성능와 약한 상관관계를 보이며, 일반화 능력의 열악한 대체 지표임을 시사한다.
  • 결과적으로 모델은 비슷한 쿼리에서는 잘 보간할 수 있지만, 진정으로 새로운 쿼리에 대해 일반화하지 못함을 드러내며, 현재 평가 관행에 심각한 격차가 있음을 입증한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.