Skip to main content
QUICK REVIEW

[論文レビュー] Semi-Supervised Quantile Estimation: Robust and Efficient Inference in High Dimensional Settings

Abhishek Chakrabortty, Guorong Dai|arXiv (Cornell University)|Jan 25, 2022
Statistical Methods and Inference被引用数 4
ひとこと要約

本稿では、高次元設定における応答分位数の推定精度と推論効率を向上させるために、大規模なラベルなしデータセットを活用する半教師付き分位数推定手法を提案する。モデルに依存しない補完とデバイアス補正ステップ、および1ステップ更新を組み合わせることで、補完モデルが誤りであっても根n一致性と漸近正規性を達成し、モデルが正しく指定されている場合には半パラメトリック効率性を達成する。

ABSTRACT

We consider quantile estimation in a semi-supervised setting, characterized by two available data sets: (i) a small or moderate sized labeled data set containing observations for a response and a set of possibly high dimensional covariates, and (ii) a much larger unlabeled data set where only the covariates are observed. We propose a family of semi-supervised estimators for the response quantile(s) based on the two data sets, to improve the estimation accuracy compared to the supervised estimator, i.e., the sample quantile from the labeled data. These estimators use a flexible imputation strategy applied to the estimating equation along with a debiasing step that allows for full robustness against misspecification of the imputation model. Further, a one-step update strategy is adopted to enable easy implementation of our method and handle the complexity from the non-linear nature of the quantile estimating equation. Under mild assumptions, our estimators are fully robust to the choice of the nuisance imputation model, in the sense of always maintaining root-n consistency and asymptotic normality, while having improved efficiency relative to the supervised estimator. They also attain semi-parametric optimality if the relation between the response and the covariates is correctly specified via the imputation model. As an illustration of estimating the nuisance imputation function, we consider kernel smoothing type estimators on lower dimensional and possibly estimated transformations of the high dimensional covariates, and we establish novel results on their uniform convergence rates in high dimensions, involving responses indexed by a function class and usage of dimension reduction techniques. These results may be of independent interest. Numerical results on both simulated and real data confirm our semi-supervised approach's improved performance, in terms of both estimation and inference.

研究の動機と目的

  • 現代のバイオメディカル研究や観察的研究において、ラベル付きデータが限られていることが原因で発生する分位数推定の統計的パワーの低さという課題に対処すること。
  • 豊富なラベルなし共変量を活用する半教師付き推論フレームワークを構築し、推定の効率性とロバスト性を向上させること。
  • 補完モデルが誤って指定されていても、分位数推定量の根n一致性と漸近正規性を保証すること。
  • モデルが正しく指定されている場合に半パラメトリック効率性を達成し、教師あり推定量よりも精度を高めること。
  • 低次元化された関数クラスに依存する応答に対して、高次元におけるカーネルスムージングのための新しい一様収束レートを確立すること。

提案手法

  • 柔軟な補完戦略を分位数推定方程式に適用することで、家族的な半教師付き推定量を提案する。
  • 補完モデルの誤りに対してロバスト性を確保するため、デバイアス補正ステップを組み込む。これにより、根n一致性と漸近正規性が維持される。
  • 非線形な分位数推定方程式の取り扱いを簡素化し、実装を容易にするために1ステップ更新を採用する。
  • 高次元共変量の低次元射影上でカーネルスムージングを用いて、ネガティブな補完関数を推定する。
  • カーネルスムージングの前段階として、次元削減技術(例:線形回帰、スライス逆回帰)を用いて有効次元を低減する。
  • 関数クラスに依存する応答を持つ高次元設定におけるカーネルスムージング推定量のための、新しい一様収束レートを確立する。

実験結果

リサーチクエスチョン

  • RQ1ラベルなしデータを活用することで、ラベル付きデータが限られた高次元設定における分位数推定の精度を向上させることができるか?
  • RQ2半教師付き分位数推定において、補完モデルの誤りに対してどのようにロバスト性を確保できるか?
  • RQ3モデルが誤って指定されている場合に、半教師付き分位数推定量の漸近的挙動はどのようになるか?
  • RQ4モデルが正しく指定されている場合に、分位数推定で半パラメトリック効率性を達成できるか?
  • RQ5次元削減された高次元共変量にカーネルスムージングを適用した場合、推定量の均一収束レートはどのようになるか?

主な発見

  • 提案手法の推定量は、いかなる補完モデルに対しても根n一致性と漸近正規性を維持し、モデルの誤りに対してもロバストであることが保証される。
  • 教師あり標本分位数と比較して、シミュレーションにおいて最大30%の相対的効率向上を達成する。
  • 補完モデルが応答の条件付き平均を正しく指定している場合、半パラメトリック効率性が達成される。
  • 次元削減された共変量におけるカーネルスムージングは、高次元入力に対して有利にスケーリングされる一様収束レートを達成する。
  • 数値結果から、半教師付きアプローチは教師あり手法と比較して、より正確な分位数推定と95%信頼区間の良好なカバレッジを実現することが示された。
  • 実データ解析において、高次元共変量を有する大規模な健康調査データセットで、本手法はより優れた推論性能を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。