[论文解读] Identifying Causes of Test Unfairness: Manipulability and Separability
该论文提出基于处理分解和可分离效应的因果框架,用以在性别、ELL 状态等不可操作的分组变量中识别可干预、可操作的测试不公源,并通过因果森林与 BART 展示其检测方法。
Differential item functioning (DIF) is a widely used statistical notion for identifying items that may disadvantage specific groups of test-takers. These groups are often defined by non-manipulable characteristics, e.g., gender, race/ethnicity, or English-language learner (ELL) status. While DIF can be framed as a causal fairness problem by treating group membership as the treatment variable, this invokes the long-standing controversy over the interpretation of causal effects for non-manipulable treatments. To better identify and interpret causal sources of DIF, this study leverages an interventionist approach using treatment decomposition proposed by Robins and Richardson (2010). Under this framework, we can decompose a non-manipulable treatment into intervening variables. For example, ELL status can be decomposed into English vocabulary unfamiliarity and classroom learning barriers, each of which influences the outcome through different causal pathways. We formally define separable DIF effects associated with these decomposed components, depending on the absence or presence of item impact, and provide causal identification strategies for each effect. We then apply the framework to biased test items in the SAT and Regents exams. We also provide formal detection methods using causal machine learning methods, namely causal forests and Bayesian additive regression trees, and demonstrate their performance through a simulation study. Finally, we discuss the implications of adopting interventionist approaches in educational testing practices.
研究动机与目标
- 推动将 DIF 解释建立在超越不可操作分组效应(如性别、ELL)的因果解释之上。
- 提出一种干预式处理分解方法,以识别 DIF 的可分离源。
- 在 FFRCISTG 框架下定义简单可分离 DIF 与广义可分离 DIF,并建立识别策略。
- 通过 SAT 与 Regents 考试题目来说明该框架,并讨论对测试实践的实际影响。
提出的方法
- 采用处理分解,将不可操作的处理拆分为通过不同因果路径起作用的干预组成部分。
- 用 SWIGs 与潜在结果定义可分离直接效应(SDE)和可分离间接效应(SIE)及其条件形式。
- 在不产生题目影响的简单可分离 DIF(无题目影响)与具有题目影响的广义可分离 DIF(含题目影响)基础上,给出相应的识别公式。
- 在一致性、忽略性、正性和可弃去组件条件(未来试验 G)下推导识别公式。
- 提出使用因果森林与贝叶斯加性回归树(BART)进行检测的方法。
- 将框架应用于 SAT 数学与 Regents 数学考试的真实题目,并通过仿真实验评估性能。
实验结果
研究问题
- RQ1来自不可操作分组处理分解的可分离因果 DIF 源是什么?
- RQ2在不存在或存在题目影响时,如何定义和识别可分离 DIF?
- RQ3因果机器学习方法(因果森林、BART)是否能在实践中有效检测可分离 DIF?
- RQ4SAT 与 Regents 的例子揭示在干预式分解下可操作的不公来源有哪些?
主要发现
- 提出将可分离 DIF 定义为在由分解处理组成部分定义的两种可干预世界之间的题目功能差异。
- 在 FFRCISTG 下提供简单和广义可分离 DIF 的非参数识别策略,并给出明确假设。
- 演示使用因果森林与 BART 的检测方法,并通过仿真实验评估。
- 将该框架应用于有偏的 SAT 与 Regents 题目,以说明可分离 DIF 组件的现实世界可解释性。
- 主张采用干预式方法来识别并纠正教育测试中的可操作性不公来源。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。