[논문 리뷰] A survey on fairness of large language models in e-commerce: progress, application, and challenge
이 종합 검토는 전자상거래 분야에서 대규모 언어 모델(Large Language Models, LLMs)의 공정성에 대해 종합적인 분석을 제공하며, 제품 리뷰, 추천, 번역, Q&A 시스템 등 응용 분야를 다루고, 훈련 데이터 및 알고리즘으로부터 유래하는 편향을 규명한다. 이 논문은 개선된 공정성 측정 지표를 제안하고 AI 개발 전 과정에서 편향 완화를 주장하며, 공정하고 투명하며 신뢰할 수 있는 전자상거래 AI 시스템을 확보하기 위해 다학제적 협력을 촉구한다.
This survey explores the fairness of large language models (LLMs) in e-commerce, examining their progress, applications, and the challenges they face. LLMs have become pivotal in the e-commerce domain, offering innovative solutions and enhancing customer experiences. This work presents a comprehensive survey on the applications and challenges of LLMs in e-commerce. The paper begins by introducing the key principles underlying the use of LLMs in e-commerce, detailing the processes of pretraining, fine-tuning, and prompting that tailor these models to specific needs. It then explores the varied applications of LLMs in e-commerce, including product reviews, where they synthesize and analyze customer feedback; product recommendations, where they leverage consumer data to suggest relevant items; product information translation, enhancing global accessibility; and product question and answer sections, where they automate customer support. The paper critically addresses the fairness challenges in e-commerce, highlighting how biases in training data and algorithms can lead to unfair outcomes, such as reinforcing stereotypes or discriminating against certain groups. These issues not only undermine consumer trust, but also raise ethical and legal concerns. Finally, the work outlines future research directions, emphasizing the need for more equitable and transparent LLMs in e-commerce. It advocates for ongoing efforts to mitigate biases and improve the fairness of these systems, ensuring they serve diverse global markets effectively and ethically. Through this comprehensive analysis, the survey provides a holistic view of the current landscape of LLMs in e-commerce, offering insights into their potential and limitations, and guiding future endeavors in creating fairer and more inclusive e-commerce environments.
연구 동기 및 목표
- 전자상거래 플랫폼에 적용된 대규모 언어 모델(Large Language Models, LLMs)의 공정성 현황을 검토하기 위해.
- 훈련 데이터 및 모델 아키텍처의 편향이 전자상거래 응용 분야에서 차별적인 결과를 초래하는 방식을 규명하기 위해.
- 성별, 인종, 직업 등 분야별로 LLM이 생성한 콘텐츠의 편향을 평가하기 위한 기존의 공정성 측정 지표 및 벤치마크를 평가하기 위해.
- 전자상거래에서 AI 개발 전 과정에 공정성을 통합하기 위한 프레임워크를 제안하기 위해.
- 글로벌 전자상거래 생태계에서 LLM의 보다 공정하고 투명하며 포용적인 구현을 위한 향후 연구를 이끌기 위해.
제안 방법
- 전자상거래에서 LLM 개발 원칙을 체계적으로 검토하며, 사전 훈련, 피너터닝, 프롬프트 엔지니어링을 포함한다.
- 주요 전자상거래 응용 분야인 제품 리뷰, 추천, 번역, Q&A 시스템에서의 공정성 도전 과제를 분류하고 분석한다.
- 내재적 및 외재적 공정성 측정 지표를 평가하며, BOLD(Bias in Open-Ended Language Generation) 및 대체적 정서 편향(Counterfactual Sentiment Bias, CSB)을 포함한다.
- 감성 불균형을 민감한 특성에 따라 정량화하기 위해 워샤르슈타인-1 거리(Wasserstein-1 distance)를 사용한 공식적 공정성 측정 지표를 도입한다.
- 데이터 수집부터 배포까지 AI 파이프라인의 모든 단계에 공정성 검사를 통합하는 것을 제안한다.
- 다양한 전자상거래 환경에서 공정성 인식 모델을 확장하기 위해 도메인 적응 및 다학제적 협력을 주장한다.

실험 결과
연구 질문
- RQ1훈련 데이터 및 모델 아키텍처의 편향이 전자상거래 LLM의 공정성에 어떻게 영향을 미치는가?
- RQ2제품 추천 및 고객 지원과 같은 특정 전자상거래 응용 분야에서의 주요 공정성 도전 과제는 무엇인가?
- RQ3BOLD 및 CSB와 같은 현재의 벤치마크가 인구 집단 간 공정성을 측정하는 데 얼마나 효과적인가?
- RQ4민감한 특성에 따른 공정성 측정 지표 외에도 전자상거래 환경에서 공정성을 정확히 포괄할 수 있는 측정 지표 및 평가 프레임워크는 무엇인가?
- RQ5전자상거래 LLM의 전체 AI 개발 전 과정에 공정성을 체계적으로 통합하는 방법은 무엇인가?
주요 결과
- 전자상거래에서의 LLM은 정제되지 않은 인터넷 데이터에서 유래한 사회적 편향을 유산으로 이어가며, 감성, 독성, 표현 측면에서 차별적인 결과를 초래한다.
- BOLD 벤치마크는 자연어 프롬프트를 사용하여 성별, 인종, 종교, 직업, 정치적 이념의 다섯 분야에서 대규모 편향 평가를 가능하게 한다.
- 대체적 정서 편향(Counterfactual Sentiment Bias, CSB) 측정 지표는 워샤르슈타인-1 거리를 사용하여 공정성을 정량화한다: I.F.는 대체적 쌍 간의 개인 수준 공정성을 측정하고, G.F.는 집단 수준의 감성 불균형을 평가한다.
- 공정성 측정 지표는 민감한 특성이 다른 대체적 쌍에 대해 모델이 유의미하게 다른 감성 점수를 생성하는 것으로 나타나 측정 가능한 편향이 있음을 시사한다.
- 현재의 공정성 평가 방법은 정적 벤치마크에 의존하여 실세계 배포 파이프라인에 통합되지 못하고 있다.
- 향후 발전은 표준화된 평가 프레임워크를 바탕으로 데이터 수집, 모델 훈련, 지속적 모니터링 단계에 공정성을 통합하는 데 달려 있다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.