[논문 리뷰] Identification of Bias Against People with Disabilities in Sentiment Analysis and Toxicity Detection Models
이 논문은 감성 및 독성 모델에서 장애 관련 편향을 탐지하기 위해 설계된 1,126개 문장으로 구성된 Bias Identification Test in Sentiments (BITS) 코퍼스를 소개한다. 이는 VADER, TextBlob, Google Cloud NLP, DistilBERT와 같은 널리 사용되는 네 가지 감성 분석 도구와 Toxic Comment Classification 및 Unintended Bias in Toxic Comments와 같은 두 가지 독성 분류기에서 통계적으로 유의미한 부정적 편향을 입증하며, p값이 2e-16에 이르는 등 장애 관련 용어가 일관되게 부정적이거나 독성으로 잘못 분류됨을 보여준다.
Sociodemographic biases are a common problem for natural language processing, affecting the fairness and integrity of its applications. Within sentiment analysis, these biases may undermine sentiment predictions for texts that mention personal attributes that unbiased human readers would consider neutral. Such discrimination can have great consequences in the applications of sentiment analysis both in the public and private sectors. For example, incorrect inferences in applications like online abuse and opinion analysis in social media platforms can lead to unwanted ramifications, such as wrongful censoring, towards certain populations. In this paper, we address the discrimination against people with disabilities, PWD, done by sentiment analysis and toxicity classification models. We provide an examination of sentiment and toxicity analysis models to understand in detail how they discriminate PWD. We present the Bias Identification Test in Sentiments (BITS), a corpus of 1,126 sentences designed to probe sentiment analysis models for biases in disability. We use this corpus to demonstrate statistically significant biases in four widely used sentiment analysis tools (TextBlob, VADER, Google Cloud Natural Language API and DistilBERT) and two toxicity analysis models trained to predict toxic comments on Jigsaw challenges (Toxic comment classification and Unintended Bias in Toxic comments). The results show that all exhibit strong negative biases on sentences that mention disability. We publicly release BITS Corpus for others to identify potential biases against disability in any sentiment analysis tools and also to update the corpus to be used as a test for other sociodemographic variables as well.
연구 동기 및 목표
- 장애가 있는 사람(PWD)에 대한 감성 및 독성 탐지 모델 내 암묵적 편향을 조사하고 폭 드러내는 것.
- 특히 소셜 미디어 데이터로 훈련된 모델에서 발생하는 능력주의적 편향이라는 미처 다루어지지 않은 문제를 해결하는 것.
- 초기에는 장애에 집중하여, 사회적 인구통계적 편향을 탐지하기 위한 표준화되고 모델에 종속되지 않는 테스트를 개발하는 것.
- 널리 사용되는 공개 도구에서 체계적인 편향 평가를 가능하게 하여 NLP 분야의 공정성과 포용성을 증진하는 것.
- BITS 코퍼스를 공개 자원으로 제공하여 향후 다른 사회적 인구통계적 집단으로의 확장과 지속적인 편향 탐지에 기여하는 것.
제안 방법
- 장애 관련 용어와 사회적 집단에 대한 참조를 다루는 1,126개의 영어 문장으로 구성된 감성 중립 및 감성 포함 코퍼스를 설계.
- 편향을 탐지하기 위한 다양한 문맥적 제어 문장을 생성하기 위해 템플릿 기반 접근 방식을 사용하여 BITS 코퍼스를 구축.
- VADER, TextBlob, Google Cloud Natural Language API, DistilBERT와 같은 네 가지 널리 사용되는 감성 분석 모델을 평가하기 위해 코퍼스를 적용.
- Detoxify 라이브러리의 두 가지 독성 탐지 모델인 Toxic Comment Classification 및 Unintended Bias in Toxic Comments를 테스트.
- 사회적 인구통계적 집단 지표를 사용한 선형 회귀 분석을 통해 모델 출력에서의 편향의 통계적 유의성을 측정.
- 감성 및 독성 점수를 장애가 있는 사람(PWD)을 포함한 다양한 사회적 인구통계적 집단 간 비교하기 위해 p값 분석을 수행.
실험 결과
연구 질문
- RQ1널리 사용되는 감성 분석 모델이 장애가 있는 사람에 대해 통계적으로 유의미한 편향을 보이는가?
- RQ2장애 관련 용어를 언급하는 문장을 처리할 때 독성 탐지 모델은 어떻게 성능을 보이는가?
- RQ3표준화되고 모델에 종속되지 않는 코퍼스가 NLP 시스템 내 암묵적 편향을 효과적으로 식별할 수 있는가?
- RQ4소셜 미디어 데이터로 훈련된 모델이 장애 관련 콘텐츠를 부정적이거나 독성으로 잘못 분류하는 정도는 어느 정도인가?
- RQ5BITS 코퍼스는 인종, 성별, 사회경제적 지위와 같이 장애 이외의 요소에 대한 편향 탐지로 확장될 수 있는가?
주요 결과
- VADER, TextBlob, Google Cloud NLP, DistilBERT와 같은 네 가지 감성 분석 모델 모두 장애 관련 콘텐츠에 대해 통계적으로 유의미한 부정적 편향을 보이며, p값이 2e-16에 이르는 등 매우 높은 유의성 수준을 확보한다.
- Toxic Comment Classification 및 Unintended Bias in Toxic Comments 모델 모두 '자폐성'이나 '정신적 장애'와 같은 용어에 대해 유의미하게 높은 독성 점수를 할당하여 강력한 암묵적 능력주의적 편향을 보여준다.
- DSBL:S(장애 특이적 편향) 비교의 p값은 VADER와 DistilBERT에서 모두 2e-16이었으며, 관찰된 편향의 극도로 높은 통계적 유의성을 확인한다.
- 장애 관련 용어에 대한 독성 점수의 낮은 표준편차는 잡음이 아닌 일관된 과도한 분류를 의미하며, 이는 랜덤한 오류가 아님을 시사한다.
- PWD에 대해 문맥적으로 민감한 훈련을 받은 Toxicity_Biased 모델은 원본 모델과 비교해 성능 향상이 없었으며, 이는 훈련 데이터가 여전히 왜곡되어 있는 한 편향 완화 노력이 효과가 없을 수 있음을 시사한다.
- BITS 코퍼스는 테스트된 모든 모델에서 편향을 노출하여, NLP 시스템에서 장애 관련 편향을 평가하는 재현 가능하고 공개 가능한 기준으로서의 유용성을 입증했다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.