Skip to main content
QUICK REVIEW

[Paper Review] Semi-Supervised Quantile Estimation: Robust and Efficient Inference in High Dimensional Settings

Abhishek Chakrabortty, Guorong Dai|arXiv (Cornell University)|Jan 25, 2022
Statistical Methods and Inference4 citations
TL;DR

This paper proposes a semi-supervised quantile estimation method that leverages a large unlabeled dataset to improve estimation accuracy and inference efficiency for response quantiles in high-dimensional settings. By combining model-agnostic imputation with a debiasing step and one-step update, the method achieves root-n consistency and asymptotic normality even under imputation model misspecification, while attaining semi-parametric efficiency when the model is correctly specified.

ABSTRACT

We consider quantile estimation in a semi-supervised setting, characterized by two available data sets: (i) a small or moderate sized labeled data set containing observations for a response and a set of possibly high dimensional covariates, and (ii) a much larger unlabeled data set where only the covariates are observed. We propose a family of semi-supervised estimators for the response quantile(s) based on the two data sets, to improve the estimation accuracy compared to the supervised estimator, i.e., the sample quantile from the labeled data. These estimators use a flexible imputation strategy applied to the estimating equation along with a debiasing step that allows for full robustness against misspecification of the imputation model. Further, a one-step update strategy is adopted to enable easy implementation of our method and handle the complexity from the non-linear nature of the quantile estimating equation. Under mild assumptions, our estimators are fully robust to the choice of the nuisance imputation model, in the sense of always maintaining root-n consistency and asymptotic normality, while having improved efficiency relative to the supervised estimator. They also attain semi-parametric optimality if the relation between the response and the covariates is correctly specified via the imputation model. As an illustration of estimating the nuisance imputation function, we consider kernel smoothing type estimators on lower dimensional and possibly estimated transformations of the high dimensional covariates, and we establish novel results on their uniform convergence rates in high dimensions, involving responses indexed by a function class and usage of dimension reduction techniques. These results may be of independent interest. Numerical results on both simulated and real data confirm our semi-supervised approach's improved performance, in terms of both estimation and inference.

Motivation & Objective

  • To address the challenge of low statistical power in high-dimensional quantile estimation due to limited labeled data in modern biomedical and observational studies.
  • To develop a semi-supervised inference framework that leverages abundant unlabeled covariates to improve estimation efficiency and robustness.
  • To ensure root-n consistency and asymptotic normality of quantile estimators even when the imputation model is misspecified.
  • To achieve semi-parametric efficiency under correct model specification, enhancing precision over supervised estimators.
  • To establish novel uniform convergence rates for kernel smoothing in high dimensions with dimension reduction, applicable to nuisance function estimation.

Proposed method

  • Proposes a family of semi-supervised estimators based on a flexible imputation strategy applied to the quantile estimating equation.
  • Incorporates a debiasing step to ensure robustness against imputation model misspecification, preserving root-n consistency and asymptotic normality.
  • Employs a one-step update to simplify implementation and handle the nonlinearity of the quantile estimating equation.
  • Uses kernel smoothing on lower-dimensional projections of high-dimensional covariates to estimate the nuisance imputation function.
  • Applies dimension reduction techniques (e.g., linear regression, sliced inverse regression) to reduce effective dimensionality before kernel smoothing.
  • Establishes new uniform convergence rates for kernel smoothing estimators in high-dimensional settings with function-class-indexed responses.

Experimental results

Research questions

  • RQ1Can unlabeled data be leveraged to improve quantile estimation accuracy in high-dimensional settings with limited labeled data?
  • RQ2How can robustness to imputation model misspecification be ensured in semi-supervised quantile estimation?
  • RQ3What is the asymptotic behavior of semi-supervised quantile estimators under model misspecification?
  • RQ4Can semi-parametric efficiency be achieved in quantile estimation under correct model specification?
  • RQ5What are the uniform convergence rates of kernel smoothing estimators when applied to high-dimensional, dimensionally reduced covariates?

Key findings

  • The proposed estimators maintain root-n consistency and asymptotic normality under any imputation model, ensuring robustness to model misspecification.
  • The estimators achieve improved efficiency over supervised sample quantiles, with relative efficiency gains of up to 30% in simulations.
  • Semi-parametric efficiency is attained when the imputation model correctly specifies the conditional mean of the response.
  • Kernel smoothing on dimensionally reduced covariates achieves uniform convergence rates that scale favorably with high-dimensional input.
  • Numerical results show that the semi-supervised approach yields more accurate quantile estimates and better coverage of 95% confidence intervals compared to supervised methods.
  • In real data analysis, the method demonstrated improved inference performance on a large-scale health survey dataset with high-dimensional covariates.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.