Skip to main content
QUICK REVIEW

[论文解读] Optimal Multilevel Matching in Clustered Observational Studies: A Case Study of the Effectiveness of Private Schools Under a Large-Scale Voucher System

Luke Keele, José R. Zubizarreta|arXiv (Cornell University)|Sep 30, 2014
Advanced Causal Inference Techniques参考文献 6被引用 5
一句话总结

本文提出了一种新颖的聚类观察性研究中的最优多水平匹配策略,颠覆了传统做法,先在不同聚类之间匹配个体,再匹配聚类本身。该方法结合动态规划与整数规划,无需倾向得分估计即可在个体和聚类两个层面实现协变量平衡,并将其应用于智利全国教育券制度下私立学校效果的评估,发现经过完整的协变量调整后,私立学校在考试成绩方面并无显著优势。

ABSTRACT

A distinctive feature of a clustered observational study is its multilevel or nested data structure arising from the assignment of treatment, in a non-random manner, to groups or clusters of units or individuals. Examples are ubiquitous in the health and social sciences including patients in hospitals, employees in firms, and students in schools. What is the optimal matching strategy in a clustered observational study? At first thought, one might start by matching clusters of individuals and then, within matched clusters, continue by matching individuals. But as we discuss in this paper, the optimal strategy is the opposite: in typical applications, where the intracluster correlation is not perfect, it is best to first match individuals and, once all possible combinations of matched individuals are known, then match clusters. In this paper we use dynamic and integer programming to implement this strategy and extend optimal matching methods to hierarchical and multilevel settings. Among other matched designs, our strategy can approximate a paired clustered randomized study by finding the largest sample of matched pairs of treated and control individuals within matched pairs of treated and control clusters that is balanced according to specifications given by the investigator. This strategy directly balances covariates both at the cluster and individual levels and does not require estimating the propensity score, although the propensity score can be balanced as an additional covariate. We illustrate our results with a case study of the comparative effectiveness of public versus private voucher schools in Chile, a question of intense policy debate in the country at the present.

研究动机与目标

  • 解决治疗分配在聚类层面但结果在个体层面测量的聚类观察性研究中的协变量不平衡问题。
  • 开发一种能够同时在个体和聚类两个层面平衡协变量的匹配策略,以改善多水平数据结构下的因果推断。
  • 提供一种不依赖倾向得分估计的方法,同时仍可将倾向得分作为额外的平衡变量纳入。
  • 通过智利大规模教育券制度下私立学校效果的案例研究,展示该方法的实用性。

提出的方法

  • 该方法颠覆了传统的匹配顺序:先在所有处理组与对照组聚类的组合中匹配个体,再对聚类本身进行匹配。
  • 采用动态规划与整数规划实现最优匹配,最小化总协变量距离或在平衡约束下最大化样本量。
  • 该方法支持最优匹配(最小化距离)与基数匹配(最大化平衡样本量),并具备平衡协变量整个分布的灵活性。
  • 该方法直接在两个层面平衡可观测协变量,可通过形成匹配的聚类对与个体对,近似实现配对的聚类随机化实验。
  • 该方法可扩展至三个或更多层次(如学生 within 学校 within 区域),并支持对未观测混杂因素的敏感性分析。

实验结果

研究问题

  • RQ1当治疗在聚类层面分配时,在聚类观察性研究中,最优匹配策略是什么?
  • RQ2在聚类之前先匹配个体,是否能比传统的先匹配聚类的方法在两个层面实现更好的协变量平衡?
  • RQ3能否设计一种匹配策略,在不估计倾向得分的情况下,同时在个体和聚类两个层面实现协变量平衡?
  • RQ4在现实政策情境中,例如评估全国教育券制度下私立学校的效果,该方法表现如何?
  • RQ5在智利教育券制度的案例研究中,研究结果对未观测混杂因素的稳健性如何?

主要发现

  • 所提出的方法成功在个体和聚类两个层面实现了协变量平衡,显著减少了匹配后私立与公立学校学生在城市居住率和经济社会地位方面的差异。
  • 匹配后,私立学校学生中城市居民的比例从未匹配样本中的79%上升至85%,更接近总体样本的分布。
  • 匹配样本中,私立与公立学校学生的标准化考试成绩差异极小,语言成绩分别为244.84和244.92。
  • 在完成完整的协变量调整后,分析未发现私立教育券学校在提升学生学业表现方面具有统计上显著的优势。
  • 敏感性分析表明,结果对未观测混杂因素导致的隐藏偏差具有稳健性,支持研究发现的内部有效性。
  • 该方法通过形成处理组与对照组聚类及其内部个体的匹配对,成功近似实现了一项配对的聚类随机化实验。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。