[Paper Review] Fashion Conversation Data on Instagram
This paper introduces a novel dataset of 24,752 Instagram fashion images labeled with visual categories and emotional cues, using manual tagging and a trained CNN to classify image types. It finds that 'body snap' and 'face' images—portraying fashion naturally—generate significantly more engagement than 'product-only' shots, despite being less frequent, highlighting the importance of authentic visual storytelling in fashion marketing on social media.
The fashion industry is establishing its presence on a number of visual-centric social media like Instagram. This creates an interesting clash as fashion brands that have traditionally practiced highly creative and editorialized image marketing now have to engage with people on the platform that epitomizes impromptu, realtime conversation. What kinds of fashion images do brands and individuals share and what are the types of visual features that attract likes and comments? In this research, we take both quantitative and qualitative approaches to answer these questions. We analyze visual features of fashion posts first via manual tagging and then via training on convolutional neural networks. The classified images were examined across four types of fashion brands: mega couture, small couture, designers, and high street. We find that while product-only images make up the majority of fashion conversation in terms of volume, body snaps and face images that portray fashion items more naturally tend to receive a larger number of likes and comments by the audience. Our findings bring insights into building an automated tool for classifying or generating influential fashion information. We make our novel dataset of {24,752} labeled images on fashion conversations, containing visual and textual cues, available for the research community.
Motivation & Objective
- To understand how fashion brands and individuals use visual content on Instagram to drive engagement.
- To identify visual features—such as image type, facial expressions, and branding—that correlate with higher likes and comments.
- To build a labeled dataset to enable future research on fashion trend detection, personalized recommendations, and automated content analysis.
- To address the lack of publicly available, fine-grained fashion image datasets with visual and textual annotations.
- To explore the gap between high-volume fashion posts and those that actually generate audience interaction.
Proposed method
- Manual annotation of 24,752 Instagram fashion images into five visual categories: selfie, body snap, marketing shot, product-only, and non-fashion.
- Training a convolutional neural network (CNN) on the annotated data to automatically classify image categories and detect visual features like 'face' and 'brand logo'.
- Applying a pre-trained facial emotion classifier to detect emotional cues (e.g., happiness, neutral) in images containing faces.
- Using regression and ANOVA tests to analyze the relationship between visual features and audience engagement (likes/comments).
- Validating the CNN model’s accuracy through quantitative evaluation on a held-out test set.
- Employing hashtag-based data collection to gather posts from 48 fashion brands across four types: mega couture, small couture, designers, and high street.
Experimental results
Research questions
- RQ1Which visual image types in fashion posts on Instagram are most effective in generating likes and comments?
- RQ2How do facial expressions and emotional cues in fashion images influence audience engagement?
- RQ3What is the relationship between image category (e.g., product-only vs. body snap) and the volume of user engagement?
- RQ4How do different fashion brand types (e.g., high street vs. couture) differ in their visual content strategies on Instagram?
- RQ5What challenges arise in classifying non-canonical fashion images such as clickbait or zoomed-in textile shots?
Key findings
- Despite being the most common type, 'product-only' images accounted for the highest volume but received the least engagement, with only 19% of likes despite making up the majority of posts.
- Body snap and selfie images—featuring fashion items in natural, contextual settings—received 53% of total likes, despite constituting only 31% of the total posts.
- Images containing faces, especially with expressions of happiness or neutrality, showed a statistically significant positive correlation with higher like counts.
- The CNN model achieved high accuracy in classifying the five main visual categories, demonstrating the feasibility of automated fashion image classification.
- Non-canonical image types such as advertisements, clickbait, and zoomed-in textile shots were frequently present but did not fit into the primary categories, revealing limitations in current classification schemes.
- The dataset revealed that some fashion hashtags were used in non-fashion contexts, indicating the presence of image spam and the need for better signal detection in user-generated content.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.