Skip to main content
QUICK REVIEW

[Paper Review] Methods and advancement of content-based fashion image retrieval: A Review

Amin Muhammad Shoib, Summaira Jabeen|arXiv (Cornell University)|Mar 30, 2023
Aesthetic Perception and Analysis4 citations
TL;DR

This paper presents a comprehensive review of content-based fashion image retrieval (CBFIR) methods, categorizing them into image-guided, image+text-guided, sketch-guided, and video-guided approaches. It analyzes recent advancements (2017–2022), evaluates network architectures, loss functions, datasets, and metrics, and identifies key challenges and future research directions in fashion retrieval systems.

ABSTRACT

Content-based fashion image retrieval (CBFIR) has been widely used in our daily life for searching fashion images or items from online platforms. In e-commerce purchasing, the CBFIR system can retrieve fashion items or products with the same or comparable features when a consumer uploads a reference image, image with text, sketch or visual stream from their daily life. This lowers the CBFIR system reliance on text and allows for a more accurate and direct searching of the desired fashion product. Considering recent developments, CBFIR still has limits when it comes to visual searching in the real world due to the simultaneous availability of multiple fashion items, occlusion of fashion products, and shape deformation. This paper focuses on CBFIR methods with the guidance of images, images with text, sketches, and videos. Accordingly, we categorized CBFIR methods into four main categories, i.e., image-guided CBFIR (with the addition of attributes and styles), image and text-guided, sketch-guided, and video-guided CBFIR methods. The baseline methodologies have been thoroughly analyzed, and the most recent developments in CBFIR over the past six years (2017 to 2022) have been thoroughly examined. Finally, key issues are highlighted for CBFIR with promising directions for future research.

Motivation & Objective

  • To provide a systematic review of recent advancements in content-based fashion image retrieval (CBFIR) from 2017 to 2022.
  • To categorize CBFIR methods into four distinct modalities: image-guided, image+text-guided, sketch-guided, and video-guided retrieval.
  • To analyze and compare the performance of CBFIR models based on network architectures, loss functions, datasets, and evaluation metrics.
  • To identify persistent challenges such as occlusion, viewpoint variation, and limited annotated data in real-world fashion retrieval.
  • To propose future research directions, including improved sketch datasets, integration of consumer purchase patterns, and enhanced video-to-shop retrieval systems.

Proposed method

  • Categorization of CBFIR methods into four main groups: image-guided, image+text-guided, sketch-guided, and video-guided retrieval.
  • Systematic analysis of deep learning-based architectures used in CBFIR, including CNNs, attention mechanisms, and multimodal fusion networks.
  • Evaluation of loss functions such as triplet loss, contrastive loss, and cross-entropy for feature embedding and similarity learning.
  • Review and comparison of benchmark datasets including DeepFashion, FashionIQ, and sketch-specific collections like Sketchy-200K.
  • Incorporation of attribute prediction and style modeling to enhance image-guided retrieval performance.
  • Integration of multimodal signals (text, audio, user feedback) in video-guided and interactive retrieval systems to improve contextual understanding.

Experimental results

Research questions

  • RQ1What are the key methodological advancements in image-guided CBFIR over the past six years?
  • RQ2How do multimodal (image + text) CBFIR methods improve retrieval accuracy compared to unimodal approaches?
  • RQ3What are the primary challenges in sketch-guided CBFIR, and how can they be addressed through improved datasets and segmentation?
  • RQ4What are the main limitations in video-guided CBFIR, particularly regarding viewpoint variation and sparse annotation?
  • RQ5What future research directions can enhance the robustness and accuracy of CBFIR systems in real-world e-commerce environments?

Key findings

  • Image-guided CBFIR methods achieve high accuracy using deep CNNs and attention mechanisms, but performance degrades under occlusion and viewpoint changes.
  • Image+text-guided CBFIR models show improved retrieval performance by leveraging multimodal embeddings, especially when text attributes are well-annotated.
  • Sketch-guided CBFIR remains challenging due to style variability and lack of texture/color information, with performance heavily dependent on the quality of sketch datasets.
  • Video-guided CBFIR is limited by insufficient publicly available video-to-shop datasets and the high visual variability caused by dynamic viewpoints and occlusions.
  • Current CBFIR systems struggle with cross-domain retrieval, particularly when reference images differ significantly in style, lighting, or pose from database items.
  • Future improvements are expected through the integration of real-time consumer purchase patterns and enhanced data augmentation for sketch and video modalities.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.