Skip to main content
QUICK REVIEW

[論文レビュー] Doubly Robust Calibration of Prediction Sets under Covariate Shift

Yachong Yang, Arun Kumar Kuchibhotla|arXiv (Cornell University)|Mar 3, 2022
Advanced Causal Inference Techniques被引用数 6
ひとこと要約

本稿は、有効な影響関数と機械学習予測子を活用して、共変量シフト下での良好にキャリブレーションされた予測集合を構築する二重ロバスト枠組みを提案する。被験者割合の推定または条件付き応答分布の推定が一貫している限り、漸近的に有効なカバレッジを保証する。これは、モデルの誤指定に対してロバストである非交換可能なデータへの順応的予測の拡張である。

ABSTRACT

Conformal prediction has received tremendous attention in recent years and has offered new solutions to problems in missing data and causal inference; yet these advances have not leveraged modern semiparametric efficiency theory for more robust and efficient uncertainty quantification. In this paper, we consider the problem of obtaining distribution-free prediction regions accounting for a shift in the distribution of the covariates between the training and test data. Under an explainable covariate shift assumption analogous to the standard missing at random assumption, we propose three variants of a general framework to construct well-calibrated prediction regions for the unobserved outcome in the test sample. Our approach is based on the efficient influence function for the quantile of the unobserved outcome in the test population combined with an arbitrary machine learning prediction algorithm, without compromising asymptotic coverage. Next, we extend our approach to account for departure from the explainable covariate shift assumption in a semiparametric sensitivity analysis for potential latent covariate shift. In all cases, we establish that the resulting prediction sets eventually attain nominal average coverage in large samples. This guarantee is a consequence of the product bias form of our proposal which implies correct coverage if either the propensity score or the conditional distribution of the response is estimated sufficiently well. Our results also provide a framework for construction of doubly robust prediction sets of individual treatment effects, under both unconfoundedness and allowing for some degree of unmeasured confounding. Finally, we discuss aggregation of prediction sets from different machine learning algorithms for optimal prediction and illustrate the performance of our methods in both synthetic and real data.

研究の動機と目的

  • 訓練データとテストデータの共変量分布が異なる(共変量シフト)状況下で、良好にキャリブレーションされた予測集合を構築する課題に対処すること。
  • 2つのノイズモデル(被験者割合または条件付き応答)のいずれかが誤って指定されても、名目上のカバレッジを維持する方法を開発すること。
  • 半パラメトリック効率理論と二重ロバスト推定を組み合わせることで、順応的予測を非交換可能なデータに拡張すること。
  • 同一の理論的基盤を活用して、無作為化の下での個別的治療効果の予測集合を構築する枠組みを提供すること。
  • 尤度比パrameterizationを用いて、説明可能な共変量シフト仮定からの逸脱に対する感度分析を可能にすること。

提案手法

  • 3つのバリエーション(分割型、全量型、効率的二重ロバスト予測)を提案。それぞれが機械学習と二重ロバスト推定を組み合わせる。
  • テスト集団におけるアウトカムの分位数のための有効な影響関数を用いて予測集合を構築する。
  • 積型バイアス構造を採用:被験者割合または条件付き応答モデルのいずれかが一貫して推定されれば、カバレッジ誤差は消失する。
  • 説明可能な共変量シフト仮定の下で、テストと訓練の共変量密度比に基づく重み付けスキームを採用する。
  • Yang と Kuchibhotla (2021) が提唱した EFCP(有効な順応的予測)アルゴリズムを改変し、実験的性能を向上させる。
  • 複数の機械学習アルゴリズムからの予測集合の集約をサポートし、より高いロバスト性と精度を実現する。

実験結果

リサーチクエスチョン

  • RQ1訓練データとテストデータの共変量分布が異なる状況下でも、予測集合が名目上のカバレッジを維持できるか?
  • RQ2提案手法がカバレッジにおいて二重ロバスト性を達成するか。つまり、被験者割合または条件付き応答モデルのいずれかが一貫して推定されれば、正しいカバレッジが保証されるか?
  • RQ3この枠組みは、無作為化の下で個別的治療効果の予測集合を構築するためにどのように拡張できるか?
  • RQ4未測定の交絡要因や説明可能な共変量シフト仮定の破綻に対する感度分析に、この手法をどのように応用できるか?
  • RQ5最適な重み付けと影響関数理論を用いることで、標準的な交差検証手法よりも効率的になるか?

主な発見

  • 提案手法の予測集合は、最小限の正則性条件のもとで漸近的に有効なカバレッジを達成し、バイアスが積型の形を取る。
  • 被験者割合または条件付き応答分布のいずれかが一貫して推定されれば、カバレッジが保証される。これにより二重ロバスト性が実現される。
  • 効率的二重ロバスト予測アルゴリズムは、最適な予測区間を事前に知っているオラクルに匹敵するほどの近似的な効率性を有すると推定される。
  • この枠組みは、無作為化の下で個別的治療効果の予測に自然に拡張可能であり、個別レベルでの不確実性の定量化を可能にする。
  • 尤度比パrameterizationを用いることで感度分析を組み込み、未測定の交絡要因に対するロバスト性の評価が可能になる。
  • 合成データおよび実データにおける実験結果から、特に有限標本においてベースライン手法を上回る性能が得られた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。