[Paper Review] A survey on fairness of large language models in e-commerce: progress, application, and challenge
This survey provides a comprehensive analysis of fairness in large language models (LLMs) within e-commerce, examining their applications in product reviews, recommendations, translation, and Q&A systems, while identifying biases from training data and algorithms. It proposes improved fairness metrics, advocates for bias mitigation throughout the AI lifecycle, and calls for interdisciplinary collaboration to ensure equitable, transparent, and trustworthy e-commerce AI systems.
This survey explores the fairness of large language models (LLMs) in e-commerce, examining their progress, applications, and the challenges they face. LLMs have become pivotal in the e-commerce domain, offering innovative solutions and enhancing customer experiences. This work presents a comprehensive survey on the applications and challenges of LLMs in e-commerce. The paper begins by introducing the key principles underlying the use of LLMs in e-commerce, detailing the processes of pretraining, fine-tuning, and prompting that tailor these models to specific needs. It then explores the varied applications of LLMs in e-commerce, including product reviews, where they synthesize and analyze customer feedback; product recommendations, where they leverage consumer data to suggest relevant items; product information translation, enhancing global accessibility; and product question and answer sections, where they automate customer support. The paper critically addresses the fairness challenges in e-commerce, highlighting how biases in training data and algorithms can lead to unfair outcomes, such as reinforcing stereotypes or discriminating against certain groups. These issues not only undermine consumer trust, but also raise ethical and legal concerns. Finally, the work outlines future research directions, emphasizing the need for more equitable and transparent LLMs in e-commerce. It advocates for ongoing efforts to mitigate biases and improve the fairness of these systems, ensuring they serve diverse global markets effectively and ethically. Through this comprehensive analysis, the survey provides a holistic view of the current landscape of LLMs in e-commerce, offering insights into their potential and limitations, and guiding future endeavors in creating fairer and more inclusive e-commerce environments.
Motivation & Objective
- To examine the current state of fairness in large language models (LLMs) applied to e-commerce platforms.
- To identify how biases in training data and model architectures lead to discriminatory outcomes in e-commerce applications.
- To evaluate existing fairness metrics and benchmarks for assessing bias in LLM-generated content across domains like gender, race, and profession.
- To propose a framework for integrating fairness into the full AI development lifecycle in e-commerce.
- To guide future research toward more equitable, transparent, and inclusive deployment of LLMs in global e-commerce ecosystems.
Proposed method
- Systematically reviews the principles of LLM development in e-commerce, including pretraining, fine-tuning, and prompt engineering.
- Classifies and analyzes fairness challenges across key e-commerce applications: product reviews, recommendations, translation, and Q&A systems.
- Evaluates intrinsic and extrinsic fairness metrics, including BOLD (Bias in Open-Ended Language Generation) and Counterfactual Sentiment Bias (CSB).
- Introduces formal fairness metrics using Wasserstein-1 distance to quantify sentiment disparities across sensitive attributes.
- Proposes integrating fairness checks at every stage of the AI pipeline—from data collection to deployment.
- Advocates for domain adaptation and interdisciplinary collaboration to scale fairness-aware models across diverse e-commerce contexts.

Experimental results
Research questions
- RQ1How do biases in training data and model architectures affect fairness in e-commerce LLMs?
- RQ2What are the key fairness challenges in specific e-commerce applications such as product recommendations and customer support?
- RQ3How effective are current benchmarks like BOLD and CSB in measuring fairness across demographic groups?
- RQ4What metrics and evaluation frameworks can accurately capture fairness beyond demographic parity in e-commerce settings?
- RQ5How can fairness be systematically embedded across the entire AI development lifecycle in e-commerce LLMs?
Key findings
- LLMs in e-commerce often inherit and amplify societal biases from uncurated internet data, leading to discriminatory outcomes in sentiment, toxicity, and representation.
- The BOLD benchmark enables large-scale evaluation of bias across five domains—gender, race, religion, profession, and political ideology—using natural language prompts.
- Counterfactual Sentiment Bias (CSB) metrics quantify fairness using Wasserstein-1 distance: I.F. measures individual fairness across counterfactual pairs, and G.F. assesses group-level sentiment disparities.
- Fairness metrics show that models generate significantly different sentiment scores for counterfactuals with varying sensitive attributes, indicating measurable bias.
- Current fairness evaluation remains limited by reliance on static benchmarks and lacks integration into real-world deployment pipelines.
- Future progress depends on embedding fairness into data collection, model training, and continuous monitoring, supported by standardized assessment frameworks.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.