[논문 리뷰] Constructing interval variables via faceted Rasch measurement and multitask deep learning: a hate speech application
본 논문은 faceted Rasch item response theory와 다중 작업 딥러닝을 결합하여, 서수 설문 항목과 텍스트 데이터의 편향 제거 예측을 이용해 혐오 발언의 연속 구간 측정치를 구성하는 방법을 제시한다.
We propose a general method for measuring complex variables on a continuous, interval spectrum by combining supervised deep learning with the Constructing Measures approach to faceted Rasch item response theory (IRT). We decompose the target construct, hate speech in our case, into multiple constituent components that are labeled as ordinal survey items. Those survey responses are transformed via IRT into a debiased, continuous outcome measure. Our method estimates the survey interpretation bias of the human labelers and eliminates that influence on the generated continuous measure. We further estimate the response quality of each labeler using faceted IRT, allowing responses from low-quality labelers to be removed. Our faceted Rasch scaling procedure integrates naturally with a multitask deep learning architecture for automated prediction on new data. The ratings on the theorized components of the target outcome are used as supervised, ordinal variables for the neural networks' internal concept learning. We test the use of an activation function (ordinal softmax) and loss function (ordinal cross-entropy) designed to exploit the structure of ordinal outcome variables. Our multitask architecture leads to a new form of model interpretation because each continuous prediction can be directly explained by the constituent components in the penultimate layer. We demonstrate this new method on a dataset of 50,000 social media comments sourced from YouTube, Twitter, and Reddit and labeled by 11,000 U.S.-based Amazon Mechanical Turk workers to measure a continuous spectrum from hate speech to counterspeech. We evaluate Universal Sentence Encoders, BERT, and RoBERTa as language representation models for the comment text, and compare our predictive accuracy to Google Jigsaw's Perspective API models, showing significant improvement over this standard benchmark.
연구 동기 및 목표
- 이진 라벨이 아니라 연속 구간 변수로 복잡한 사회적 구성요소를 측정하려는 동기.
- 확장 가능한 예측 프레임워크 내에서 인간 라벨링의 편향을 제거하고 라벨러 품질을 추정하는 목표.
- 편향 제거되고 해석 가능한 예측을 위해 Rasch 기반 측정과 딥러닝을 통합하려는 목표.
제안 방법
- 혐오 발언을 여덟 개의 이론화된 구성요소로 분해하고 이를 서수 설문 항목으로 라벨링한다.
- multi-item ordinal 라벨을 연속 구간 척도로 변환하기 위하여 faceted Rasch measurement theory를 사용한다.
- 공유 가중치를 가진 다중 작업 딥러닝 모델을 훈련시켜 텍스트로부터 잠재 구성요소를 예측한다.
- 대상들의 서수 구조를 활용하기 위해 ordinal softmax 활성화 및 ordinal cross-entropy 손실을 적용한다.
- 예측에 부분 점수 IRT 변환을 적용하여 타당한 값 점수를 얻는다.
- 리뷰어 라벨러 편향을 추정하고 저품질 응답을 필터링함으로써 편향 제거를 가능하게 한다.
실험 결과
연구 질문
- RQ1혐오 발언을 faceted Rasch 프레임워크와 감독 딥러닝을 결합하여 연속 스펙트럼으로 모델링할 수 있는가?
- RQ2라벨러 편향 제거가 연속 혐오 발언 점수의 정확도와 신뢰성을 향상시키는가?
- RQ3다중 작업 아키텍처가 구성요소와 정렬된 해석 가능한 연속 예측을 제공할 수 있는가?
- RQ4서수 활성화 및 손실 함수가 표준 접근법에 비해 서수 타깃에 대한 예측을 향상시키는가?
- RQ5이 방법이 Google Jigsaw의 Perspective API와 같은 기존 벤치마크에 비해 얼마나 성능을 보이는가?
주요 결과
- 본 방법은 여덟 개의 이론화된 수준과 32–48개의 라벨링 항목에서 파생된 연속 혐오 발언 척도를 산출한다.
- YouTube, Twitter, Reddit 전역에서 10,000명의 크라우드워커가 라벨링한 50,000개 댓글 데이터셋이 구성되었다.
- 해당 접근법은 이 작업의 예측 정확도에서 Perspective API에 비해 상당한 향상을 보인다.
- Faceted Rasch 스케일링은 라벨러와 코멘트에 대한 편향 제거 기제를 통한 불변 측정을 제공한다.
- 다중 작업 모델은 각 연속 점수를 맨 끝에서 두 번째 계층의 구성 요소와 연결함으로써 해석 가능한 예측을 제공한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.