Skip to main content
QUICK REVIEW

[논문 리뷰] Selective inference after variable selection via multiscale bootstrap

Yoshikazu Terada, Hidetoshi Shimodaira|arXiv (Cornell University)|2019. 05. 25.
Statistical Methods and Inference참고 문헌 3인용 수 6
한 줄 요약

이 논문은 회귀 분석에서 변수 선택 이후 선택 편향을 보정하기 위해 다중 척도 부트스트랩 방법을 제안한다. p-값과 신뢰구간의 선택 편향을 다루며, 구체적인 모형이 아닌 변수 포함 여부에 초점을 맞춘 더 유연한 선택 이벤트를 정의함으로써, MCP와 같은 다양한 알고리즘에 대해 유효한 추론을 가능하게 한다. 계산 비용은 전통적인 부트스트랩과 유사하다.

ABSTRACT

A general resampling approach is considered for selective inference problem after variable selection in regression analysis. Even after variable selection, it is important to know whether the selected variables are actually useful by showing $p$-values and confidence intervals of regression coefficients. In the classical approach, significance levels for the selected variables are usually computed by $t$-test but they are subject to selection bias. In order to adjust the bias in this post-selection inference, most existing studies of selective inference consider the specific variable selection algorithm such as Lasso for which the selection event can be explicitly represented as a simple region in the space of the response variable. Thus, the existing approach cannot handle more complicated algorithm such as MCP (minimax concave penalty). Moreover, most existing approaches set an event, that a specific model is selected, as the selection event. This selection event is too restrictive and may reduce the statistical power, because the hypothesis selection with a specific variable only depends on whether the variable is selected or not. In this study, we consider more appropriate selection event such that the variable is selected, and propose a new bootstrap method to compute an approximately unbiased selective $p$-value for the selected variable. Our method is applicable to a wide class of variable selection algorithms. In addition, the computational cost of our method is the same order as the classical bootstrap method. Through the numerical experiments, we show the usefulness of our selective inference approach.

연구 동기 및 목표

  • 회귀 분석에서 변수 선택 이후 p-값과 신뢰구간의 선택 편향을 해결한다.
  • Lasso와 같은 특정 알고리즘에 의존하거나 제약 조건이 있는 선택 이벤트에 의존하는 기존 선택 편향 추론 방법의 한계를 극복한다.
  • MCP(최소 최대 페널티)와 같은 복잡한 변수 선택 알고리즘에 일반화 가능하게 적용 가능한 방법을 개발한다.
  • 정확한 모형 선택이 아닌 변수 포함에 초점을 맞춘 덜 제약적인 선택 이벤트를 통해 통계적 검정력을 향상시킨다.
  • 전통적인 부트스트랩과 동일한 계산 복잡도를 유지하면서도 추론의 타당성을 확보한다.

제안 방법

  • 선택 이벤트를 특정 모형의 선택이 아닌 특정 변수의 포함으로 정의한다.
  • 선택 이벤트를 조건으로 한 검정 통계량의 조건부 분포를 근사하기 위해 다중 척도 부트스트랩 프레임워크를 제안한다.
  • 재표본 추출을 통해 선택 편향 추론 프레임워크 하에서 검정 통계량의 귀무분포를 추정한다.
  • 부트스트랩 기반 추정을 통해 관측된 선택 이벤트를 조건으로 하여 근사적으로 편향이 없는 p-값을 구성한다.
  • 전통적 부트스트랩과 동일한 복잡도 수준을 유지함으로써 계산 효율성을 확보한다.
  • 비볼록 페널티(예: MCP)를 포함한 광범위한 변수 선택 알고리즘에 적용한다.

실험 결과

연구 질문

  • RQ1선택 이벤트가 특정 모형이 아니라 변수 포함일 경우, 변수 선택 이후 선택 편향을 신뢰성 있게 다룰 수 있는 방법은 무엇인가?
  • RQ2부트스트랩 기반 방법이 다양한 변수 선택 알고리즘에 걸쳐 선택 편향을 보정하는 유효한 p-값과 신뢰구간을 제공할 수 있는가?
  • RQ3이러한 방법의 계산 비용은 전통적 부트스트랩과 비교해 어떻게 되며, 동일한 수준을 유지할 수 있는가?
  • RQ4기존 방법과 비교해 제안된 방법의 통계적 검정력과 유형 I 오류 통제 능력은 어떠한가?
  • RQ5선택 이브트가 복잡하고 쉽게 기술하기 어려운 비볼록 페널티(예: MCP)로 확장 가능할 수 있는가?

주요 결과

  • 제안된 다중 척도 부트스트랩 방법은 변수 선택 이후 추론에서 선택 편향을 보정하는 근사적으로 편향 없는 p-값을 생성한다.
  • 이 방법은 Lasso와 유사한 기존 선택 편향 추론 방법으로는 다루기 어려운 MCP와 같은 광범위한 변수 선택 알고리즘에 적용 가능하다.
  • 제안된 방법의 계산 비용은 전통적 부트스트랩과 동일한 순서이므로 확장성과 실용성이 뛰어나다.
  • 수치 실험을 통해 방법이 적절한 유형 I 오류 비율을 유지하고, 난이도 높은 후행 선택 추론보다 통계적 검정력을 향상시킨다.
  • 변수 포함 기반 선택 이벤트는 모형 특화 선택 이벤트보다 더 높은 검정력을 제공한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.