Skip to main content
QUICK REVIEW

[Paper Review] Information Retrieval: Recent Advances and Beyond

Kailash Hambarde, Hugo Proença|arXiv (Cornell University)|Jan 20, 2023
Neural Networks and ApplicationsComputer Science207 references8 citations
TL;DR

A survey of information retrieval models across retrieval and ranking stages, covering term-based, semantic, and neural approaches, plus learning techniques and recent advances.

ABSTRACT

In this paper, we provide a detailed overview of the models used for information retrieval in the first and second stages of the typical processing chain. We discuss the current state-of-the-art models, including methods based on terms, semantic retrieval, and neural. Additionally, we delve into the key topics related to the learning process of these models. This way, this survey offers a comprehensive understanding of the field and is of interest for for researchers and practitioners entering/working in the information retrieval domain.

Motivation & Objective

  • Provide a comprehensive overview of IR models used in the first-stage retrieval and second-stage ranking.
  • Discuss state-of-the-art term-based, semantic, and neural approaches in IR.
  • Explain key learning paradigms and training techniques for IR models.
  • Highlight recent trends such as expansion, dense representations, and knowledge integration.

Proposed method

  • Review and synthesize literature on traditional term-based models (e.g., BM25, TF-IDF) and probabilistic/language-model approaches.
  • Summarize pioneering and contemporary semantic and lexical dependency methods.
  • Discuss deep learning architectures for semantic retrieval, including CNNs, RNNs, transformers, and pre-trained language models.
  • Examine dense and sparse representation learning, expansion techniques, and knowledge integration in IR.
  • Outline pre-training, distillation, and multi-vector representations in dense retrieval contexts.

Experimental results

Research questions

  • RQ1What are the main model families used in the first-stage retrieval and second-stage ranking in IR?
  • RQ2How have neural and dense representations influenced IR effectiveness and efficiency recently?
  • RQ3What learning and training strategies are used to improve IR performance (e.g., expansion, pre-training, distillation)?
  • RQ4How do expansion and multilingual/dense methods contribute to modern IR systems?
  • RQ5What are the key challenges and future directions identified for information retrieval research?

Key findings

  • Deep learning and pre-trained models have significantly improved IR performance across semantic and neural methods.
  • Expansion techniques (Doc2Query, query expansion) and sparse/dense representations enhance retrieval effectiveness.
  • Dense retrieval methods employ techniques like negative sampling, end-to-end training, and representation decoupling to improve open-domain QA and passage ranking.
  • Pre-training and knowledge integration (e.g., knowledge graphs) are pivotal in advancing dense retrieval and cross-modal IR.
  • There is ongoing research in efficiency-focused methods such as pruning, hashing, and multi-vector representations to balance speed and accuracy.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.