Skip to main content
QUICK REVIEW

[Paper Review] Evaluating Treatment Prioritization Rules via Rank-Weighted Average Treatment Effects

Steve Yadlowsky, Scott Fleming|arXiv (Cornell University)|Nov 15, 2021
Advanced Causal Inference Techniques19 citations
TL;DR

This paper introduces Rank-Weighted Average Treatment Effect (RATE) metrics as a general, inference-ready framework for evaluating treatment prioritization rules—regardless of whether they are derived from causal effect models or risk scores. The method enables asymptotically valid hypothesis testing and demonstrates superior power in detecting treatment effect heterogeneity compared to existing metrics like Qini, especially under varying treatment effect distributions.

ABSTRACT

There are a number of available methods for selecting whom to prioritize for treatment, including ones based on treatment effect estimation, risk scoring, and hand-crafted rules. We propose rank-weighted average treatment effect (RATE) metrics as a simple and general family of metrics for comparing and testing the quality of treatment prioritization rules. RATE metrics are agnostic as to how the prioritization rules were derived, and only assess how well they identify individuals that benefit the most from treatment. We define a family of RATE estimators and prove a central limit theorem that enables asymptotically exact inference in a wide variety of randomized and observational study settings. RATE metrics subsume a number of existing metrics, including the Qini coefficient, and our analysis directly yields inference methods for these metrics. We showcase RATE in the context of a number of applications, including optimal targeting of aspirin to stroke patients.

Motivation & Objective

  • To develop a general, statistically rigorous metric for evaluating treatment prioritization rules that are agnostic to their derivation method.
  • To enable asymptotically exact inference for prioritization rule performance using a central limit theorem for RATE estimators.
  • To compare the statistical power of RATE under different weighting functions (linear vs. logarithmic) in detecting heterogeneous treatment effects.
  • To evaluate the performance of CATE-based and risk-based prioritization rules in real-world medical applications, such as aspirin use in stroke patients.
  • To provide open-source tools, including a rank_average_treatment_effect function in the grf R package, for practical deployment.

Proposed method

  • Proposes a family of RATE metrics that weight treatment effects by rank, with linear (Qini-like) and logarithmic (AUTOC-like) weighting functions.
  • Derives a central limit theorem for RATE estimators to enable asymptotically exact inference in both randomized and observational study designs.
  • Uses augmented inverse probability weighting (AIPW) and inverse probability weighting (IPW) estimators to ensure robustness to nuisance parameter estimation.
  • Applies the method to survival outcomes with censoring, using restricted mean survival time (RMST) as the outcome metric.
  • Employs half-sample bootstrap to compute 95% confidence intervals and p-values for RATE estimates.
  • Validates the method using both synthetic simulations and real-world data from the SPRINT and ACCORD-BP trials.

Experimental results

Research questions

  • RQ1How does the choice of weighting function (linear vs. logarithmic) affect the statistical power of RATE in detecting treatment effect heterogeneity?
  • RQ2Can RATE provide valid inference for prioritization rules derived from CATE models or risk scores, regardless of their origin?
  • RQ3How do CATE-based and risk-based prioritization rules compare in performance on real-world stroke treatment data?
  • RQ4What is the robustness of RATE estimates to misspecification of nuisance parameters in real-world observational settings?
  • RQ5How well do RATE estimates generalize across different clinical trial populations, such as SPRINT and ACCORD-BP?

Key findings

  • The Qini coefficient (linear weighting) yields greater statistical power when treatment effects are substantial and diffuse across the population.
  • The AUTOC metric (logarithmic weighting) achieves greater power when treatment effects are concentrated in a small, high-benefit subgroup.
  • RATE estimates remain robust across different score types—IPW, AIPW, and Oracle scores—confirming the stability of the method under model misspecification.
  • In the combined SPRINT/ACCORD-BP analysis, the Causal Survival Forest model yielded an AUTOC estimate of -5.83 (95% CI: -14.09, 2.43), with a p-value of 0.17, indicating no significant benefit in prioritization.
  • The Framingham Risk Score and ACC/AHA Pooled Cohort Equations showed negative but non-significant AUTOC estimates, suggesting limited prioritization value in this context.
  • The method enables valid inference for survival outcomes with censoring, extending RATE’s utility to common medical study designs.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.