Skip to main content
QUICK REVIEW

[논문 리뷰] Finite-sample analysis of M-estimators using self-concordance

Dmitrii M. Ostrovskii, Francis Bach|arXiv (Cornell University)|2018. 10. 16.
Statistical Methods and Inference참고 문헌 57인용 수 11
한 줄 요약

이 논문은 자기일관성(self-concordance)을 사용하여 M-추정량의 유한표본 경계를 제공하며, 카이제곱형 초과 위험 경계에 대한 임계 표본 크기가 $ O(d ⋅ d_{\text{eff}}) $로 스케일링됨을 보여준다. 여기서 $ d $ 는 매개변수 차원이고 $ d_{\text{eff}} $ 는 모형 오특정성에 기여한다. 분석은 모집단 최소화자에서 국소적 서브가우시안성 및 곡률 가정에 기반하며, 더 강한 Dikin 타원체 이웃 조건 하에서는 더 날카운 경계 $ O(\max\{d_{\text{eff}}, d\log d\}) $ 를 얻는다. 이는 가우시안 설계를 가진 로지스틱 회귀에서 검증되었다.

ABSTRACT

The classical asymptotic theory for parametric $M$-estimators guarantees that, in the limit of infinite sample size, the excess risk has a chi-square type distribution, even in the misspecified case. We demonstrate how self-concordance of the loss allows to characterize the critical sample size sufficient to guarantee a chi-square type in-probability bound for the excess risk. Specifically, we consider two classes of losses: (i) self-concordant losses in the classical sense of Nesterov and Nemirovski, i.e., whose third derivative is uniformly bounded with the $3/2$ power of the second derivative; (ii) pseudo self-concordant losses, for which the power is removed. These classes contain losses corresponding to several generalized linear models, including the logistic loss and pseudo-Huber losses. Our basic result under minimal assumptions bounds the critical sample size by $O(d \\cdot d_{\ ext{eff}}),$ where $d$ the parameter dimension and $d_{\ ext{eff}}$ the effective dimension that accounts for model misspecification. In contrast to the existing results, we only impose local assumptions that concern the population risk minimizer $\ heta_*$. Namely, we assume that the calibrated design, i.e., design scaled by the square root of the second derivative of the loss, is subgaussian at $\ heta_*$. Besides, for type-ii losses we require boundedness of a certain measure of curvature of the population risk at $\ heta_*$.Our improved result bounds the critical sample size from above as $O(\\max\\{d_{\ ext{eff}}, d \\log d\\})$ under slightly stronger assumptions. Namely, the local assumptions must hold in the neighborhood of $\ heta_*$ given by the Dikin ellipsoid of the population risk. Interestingly, we find that, for logistic regression with Gaussian design, there is no actual restriction of conditions: the subgaussian parameter and curvature measure remain near-constant over the Dikin ellipsoid. Finally, we extend some of these results to $\\ell_1$-penalized estimators in high dimensions.

연구 동기 및 목표

  • M-추정량이 점근적 근사에서 벗어나 유한표본에서 카이제곱형 초과 위험 경계를 달성할 수 있는 표본 크기의 유한표본 영역를 특성화하기 위해.
  • 손실 함수의 자기일관성 특성을 사용하여 이러한 경계에 필요한 표본 크기 조건을 수립하기 위해.
  • 모든 매개변수에 대한 국소적 조건만을 $ \theta_* $ 에서 가정함으로써 전역 부드러움 조건을 완화하기 위해.
  • 고차원 설정에서 $ \ell_1 $-패널티 추정량으로 결과를 확장하기 위해.
  • 로지스틱 회귀에서 가우시안 설계를 사용하여 이론을 검증하고, Dikin 타원체 내에서 거의 일정한 서브가우시안 및 곡률 파라미터를 보여주기 위해.

제안 방법

  • 자기일관성(세 번째 도함수의 헤시안에 대한 3/2 승으로 제한됨)과 의사자기일관성(거듭제곱 없음)의 두 유형의 손실 함수를 도입한다.
  • 두 번째 도함수의 제곱근으로 스케일된 예측자인 캘리브레이션된 예측자를 사용하여 $ \theta_* $ 에서 국소적 서브가우시안성을 정의한다. 이는 유한표본 경계의 핵심 가정이다.
  • Dikin 타원체 이웃 조건을 적용하여 경계를 강화하며, $ \theta_* $ 근처에서 국소적 서브가우시안성과 곡률 제어를 요구한다.
  • $ \psi_2 $ 와 $ \psi_1 $-노름을 사용하여 추정 방정식의 尾행동을 제어하고, 모멘트 및 꼬리 경계 분석을 가능하게 한다.
  • 자기일관성 하에서 농도 부등식을 적용하여 초과 위험 경계를 유도하며, 표본 크기를 효율적 차원 $ d_{\text{eff}} $ 와 연결한다.
  • 유사한 국소적 가정과 곡률 제어를 사용하여 $ \ell_1 $-패널티 M-추정량으로 결과를 확장한다.

실험 결과

연구 질문

  • RQ1M-추정량이 유한표본에서 카이제곱형 초과 위험 경계를 달성하기 위해 필요한 임계 표본 크기는 무엇인가?
  • RQ2손실 함수의 자기일관성과 의사자기일관성은 유한표본 수렴 속도에 어떤 영향을 미치는가?
  • RQ3예를 들어 캘리브레이션된 예측자의 서브가우시안성과 같은 $ \theta_* $ 에서의 국소적 가정은 전역 부드러움 조건을 대체할 수 있는가?
  • RQ4Dikin 타원체 이웃 조건은 전역 가정보다 더 날카운 표본 크기 경계를 이끌어내는가?
  • RQ5이러한 경계는 $ \ell_1 $-패널티를 적용한 고차원 설정에서 어떻게 행동하는가?

주요 결과

  • 카이제곱형 초과 위험 경계에 대한 임계 표본 크기는 $ O(d \cdot d_{\text{eff}}) $ 로 제한되며, 여기서 $ d_{\text{eff}} $ 는 모형 오특정성의 영향을 반영한다.
  • 더 강한 Dikin 타원체 이웃 조건 하에서는 경계가 $ O(\max\{d_{\text{eff}}, d\log d\}) $ 로 강화된다.
  • 가우시안 설계를 가진 로지스틱 회귀에서, 서브가우시안 파라미터와 곡률 측정값이 Dikin 타원체 전체에서 거의 일정하게 유지되어 국소적 가정이 검증된다.
  • 전체 부드러움 또는 곡률 조건 없이도 $ \theta_* $ 에서의 국소적 정규성만으로도 유한표본 경계를 달성한다.
  • 분석은 고차원 설정에서 $ \ell_1 $-패널티 M-추정량으로 확장되어, 이론적으로도 보장 가능한 경계를 제공한다.
  • $ \psi_1 $-노름의 사용은 잘 지정된 모형에서 $ \psi_2 $-노름 경계를 초월하여 지수 꼬리 제어를 가능하게 한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.