[论文解读] Optimal Multilevel Matching in Clustered Observational Studies: A Case Study of the School Voucher System in Chile
本文提出了一种用于聚类观测研究的最优多水平匹配策略,颠覆了传统观念,即先在聚类内匹配个体,再匹配聚类。通过动态规划与整数规划,该方法在个体与聚类两个层面均识别出最大化的平衡匹配对样本,展示了其在评估智利学校券制度中的应用,并实现了更优的协变量平衡。
A distinctive feature of a clustered observational study is its multilevel or nested data structure arising from the assignment of treatment, in a non-random manner, to groups or clusters of individuals. Examples are ubiquitous in the health and social sciences including patients in hospitals, employees in firms, and students in schools. What is the optimal matching strategy in a clustered observational study? At first thought, one might start by matching clusters of individuals and then, within matched clusters, continue by matching individuals. But, as we discuss in this paper, the optimal strategy is the opposite: first match individuals and, once all possible combinations of matched individuals are known, then match clusters. In this paper we use dynamic and integer programming to implement this strategy and extend optimal matching methods to hierarchical and multilevel settings. In particular, our method attempts to replicate a paired clustered randomized study by finding the largest sample of matched pairs of treated and control individuals within matched pairs of treated and control clusters that is balanced according to specifications given by the user. We illustrate our method on a case study of the comparative effectiveness of public versus private voucher schools in Chile, a question of intense policy debate in the country at the present.
研究动机与目标
- 为解决在组别非随机分配的分层聚类观测研究中实现最优匹配的挑战。
- 开发一种方法,使观测数据中的匹配结果能复现配对聚类随机实验的平衡性。
- 颠覆传统匹配顺序——先匹配个体,再匹配聚类——以在两个层面均实现更优的协变量平衡。
- 提供一种用户可指定的多水平匹配框架,优先保障个体与聚类两个层面的平衡性。
- 将该方法应用于一个现实政策问题:评估智利公立与私立券校的相对有效性。
提出的方法
- 该方法使用动态规划识别聚类内所有可能的匹配个体组合。
- 随后应用整数规划,选择最优的匹配聚类集合,以最大化平衡匹配对的数量。
- 该方法确保根据用户指定的标准,实现个体层面与聚类层面协变量的平衡。
- 匹配过程按逆序进行:先匹配个体,再匹配包含这些已匹配个体的聚类。
- 该算法在保持平衡的前提下,寻求在两个层面均能获得最大可能的匹配处理组与对照组对样本。
- 该方法旨在模拟观测环境中配对聚类随机实验的设计。
实验结果
研究问题
- RQ1在多水平聚类观测研究中,最优匹配的顺序应为:先匹配个体还是先匹配聚类?
- RQ2如何将最优匹配扩展至分层数据结构,同时在多个层面保持平衡?
- RQ3在聚类观测研究中,匹配样本能否实现与配对聚类随机实验相当的平衡性?
- RQ4匹配顺序对多水平设置中协变量平衡性及因果推断有效性有何影响?
- RQ5如何利用整数规划与动态规划在实践中实现最优多水平匹配?
主要发现
- 最优匹配策略颠覆了传统做法:先匹配个体再匹配聚类,可实现更优的平衡性与更大的匹配样本。
- 该方法成功识别出在个体与聚类两个层面均具有较大规模且高度平衡的处理组与对照组匹配对样本。
- 该方法实现的平衡性可与配对聚类随机实验相媲美,从而增强因果推断的有效性。
- 在智利学校券制度中的应用表明,该方法在现实政策评估中具有可行性与实用性。
- 通过动态规划与整数规划的结合,可在用户定义的平衡约束下高效计算最优多水平匹配。
- 本研究表明,多水平匹配可显著改善嵌套数据结构观测研究中的协变量平衡性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。