[Paper Review] Fashion-Gen: The Generative Fashion Dataset and Challenge
Introduces a large, high-definition fashion image dataset paired with expert-descriptions and presents baseline results for high-resolution image generation and text-to-image synthesis, along with a community challenge.
We introduce a new dataset of 293,008 high definition (1360 x 1360 pixels) fashion images paired with item descriptions provided by professional stylists. Each item is photographed from a variety of angles. We provide baseline results on 1) high-resolution image generation, and 2) image generation conditioned on the given text descriptions. We invite the community to improve upon these baselines. In this paper, we also outline the details of a challenge that we are launching based upon this dataset.
Motivation & Objective
- Provide a large-scale, high-quality dataset of fashion images with professional descriptions and metadata.
- Enable research on text-to-image synthesis conditioned on detailed fashion descriptions.
- Offer baselines for high-resolution image generation and text-conditioned generation.
- Launch a competitive challenge to advance text-to-image generation in fashion.
Proposed method
- Dataset collection with 293,008 HD (1360x1360) fashion images from multiple angles.
- Descriptions provided by professional designers for each item.
- Baseline experiments using progressive growing GANs to generate high-res images.
- Text-to-image synthesis experiments using StackGAN-v1 and StackGAN-v2 with various text encoders.
- Pre-trained text encoders (bi-LSTM, Transformer) evaluated for alignment between descriptions and visuals.
Experimental results
Research questions
- RQ1Can high-resolution fashion images be generated realistically from text descriptions and from noise alone on a large, expert-annotated dataset?
- RQ2How do different text encoding strategies affect quality and fidelity in text-to-image synthesis for fashion items?
- RQ3What is the impact of multi-angle photography and rich metadata on generation performance?
- RQ4How do StackGAN-v1, StackGAN-v2, and progressive GANs compare on this Fashion-Gen dataset in terms of visual quality and category fidelity?
Key findings
- Progressive GANs generate 1024x1024 fashion images with high global coherence on Fashion-Gen.
- Inception scores for real data at 256x256 are higher than StackGAN-V1, StackGAN-V2, and P-GAN baselines, with StackGAN-V1 outperforming StackGAN-V2 in score but StackGAN-V2 offering better visual quality in some cases.
- Pre-training and fixing a bi-LSTM text encoder yielded better text-to-image results than other encoders tested.
- StackGAN-v1 achieved higher Inception Scores than StackGAN-v2, but StackGAN-v2 produced higher-quality images with mode-collapse challenges observed.
- Descriptive text embeddings significantly influence the quality and fidelity of generated fashion images.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.