[论文解读] Constructing interval variables via faceted Rasch measurement and multitask deep learning: a hate speech application
本文提出了一种通过将分面 Rasch 项目反应理论与多任务深度学习相结合,结合序数调查项和来自文本数据的去偏预测,来构建仇恨言论的连续区间测量的方法。
We propose a general method for measuring complex variables on a continuous, interval spectrum by combining supervised deep learning with the Constructing Measures approach to faceted Rasch item response theory (IRT). We decompose the target construct, hate speech in our case, into multiple constituent components that are labeled as ordinal survey items. Those survey responses are transformed via IRT into a debiased, continuous outcome measure. Our method estimates the survey interpretation bias of the human labelers and eliminates that influence on the generated continuous measure. We further estimate the response quality of each labeler using faceted IRT, allowing responses from low-quality labelers to be removed. Our faceted Rasch scaling procedure integrates naturally with a multitask deep learning architecture for automated prediction on new data. The ratings on the theorized components of the target outcome are used as supervised, ordinal variables for the neural networks' internal concept learning. We test the use of an activation function (ordinal softmax) and loss function (ordinal cross-entropy) designed to exploit the structure of ordinal outcome variables. Our multitask architecture leads to a new form of model interpretation because each continuous prediction can be directly explained by the constituent components in the penultimate layer. We demonstrate this new method on a dataset of 50,000 social media comments sourced from YouTube, Twitter, and Reddit and labeled by 11,000 U.S.-based Amazon Mechanical Turk workers to measure a continuous spectrum from hate speech to counterspeech. We evaluate Universal Sentence Encoders, BERT, and RoBERTa as language representation models for the comment text, and compare our predictive accuracy to Google Jigsaw's Perspective API models, showing significant improvement over this standard benchmark.
研究动机与目标
- 以连续区间变量而非二元标签来衡量复杂社会构念的动机。
- 在可扩展的预测框架内去偏人类标注并估计标注者质量的目标。
- 将基于 Rasch 的测量与深度学习整合,以实现去偏、可解释的预测。
提出的方法
- 将仇恨言论分解为八个理论化组成部分,并用序数调查项对其进行标注。
- 使用分面 Rasch 测量理论将多项序数标签转换为连续区间尺度。
- 训练一个具有共享权重的多任务深度学习模型,从文本中预测潜在组成。
- 采用序数 softmax 激活和序数交叉熵损失,利用目标的序数结构。
- 对预测结果应用部分积分 IRT 转换以获得可行的值分数。
- 通过估计评审标注者偏差并筛选低质量回答来实现去偏。
实验结果
研究问题
- RQ1在结合监督深度学习的前提下,是否可以使用分面 Rasch 框架将仇恨言论建模为连续光谱?
- RQ2去偏标注者偏差是否会提高连续仇恨言论分数的准确性和可靠性?
- RQ3多任务架构是否能够提供与组成成分相一致的可解释的连续预测?
- RQ4序数激活和损失函数相比标准方法是否能改善对序数目标的预测?
- RQ5与 Google Jigsaw 的 Perspective API 等现有基准相比,该方法的性能如何?
主要发现
- 该方法产生了一个基于八个理论等级和 32–48 个标注项的连续仇恨言论尺度。
- 构建了一个 50,000 条评论的数据集,由 10,000 名众包工作者在 YouTube、Twitter 和 Reddit 上标注。
- 在该任务的预测准确性方面,该方法显著优于 Perspective API。
- 分面 Rasch 标度提供了不变量测量,并对标注者和评论实施去偏化机制。
- 多任务模型通过在倒数第二层将每个连续分数与其组成成分联系起来,提供可解释的预测。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。