Skip to main content
QUICK REVIEW

[논문 리뷰] Robust Validation: Confident Predictions Even When Distributions Shift

Maxime Cauchois, Suyash Gupta|arXiv (Cornell University)|2020. 08. 10.
Adversarial Robustness in Machine Learning참고 문헌 50인용 수 5
한 줄 요약

이 논문은 f-divergence 구를 사용하여 가능한 분포 변화를 모델링함으로써 분포 이탈 하에서도 보장된 커버리지로 예측 집합을 생성하는 콫포멀 추론 방법을 제안한다. 교환 가능성 하에서 유한 표본 유효성을 보장하며, 이격 정도를 추정하는 방법을 도입하여 훈련 데이터와 다를 수 있는 테스트 데이터에서도 강건한 신뢰도 캘리브레이션을 달성한다.

ABSTRACT

While the traditional viewpoint in machine learning and statistics assumes training and testing samples come from the same population, practice belies this fiction. One strategy -- coming from robust statistics and optimization -- is thus to build a model robust to distributional perturbations. In this paper, we take a different approach to describe procedures for robust predictive inference, where a model provides uncertainty estimates on its predictions rather than point predictions. We present a method that produces prediction sets (almost exactly) giving the right coverage level for any test distribution in an $f$-divergence ball around the training population. The method, based on conformal inference, achieves (nearly) valid coverage in finite samples, under only the condition that the training data be exchangeable. An essential component of our methodology is to estimate the amount of expected future data shift and build robustness to it; we develop estimators and prove their consistency for protection and validity of uncertainty estimates under shifts. By experimenting on several large-scale benchmark datasets, including Recht et al.'s CIFAR-v4 and ImageNet-V2 datasets, we provide complementary empirical results that highlight the importance of robust predictive validity.

연구 동기 및 목표

  • 훈련 데이터와 다를 수 있는 테스트 데이터가 존재하는 분포 이탈 하에서도 유효한 커버리지를 유지하는 예측 추론 방법을 개발하는 것.
  • 훈련 분포의 f-divergence 기반 변형에 대해 강건한 불확실성 집합을 체계화하는 것.
  • 실제로 예상되는 데이터 이격의 크기를 추정하여 사전에 이격에 대한 지식이 없더라도 실용적인 강건성을 달성하는 것.
  • 유한 표본 유효성 보장을 보장하는 조건으로 훈련 데이터의 교환 가능성만을 가정함으로써 강력한 파라미터 가정을 피하는 것.
  • 실제 벤치마크인 CIFAR-v4와 ImageNet-V2에서의 실증적 검증을 통해 자연 분포 이탈에 대한 강건성을 입증하는 것.

제안 방법

  • 이 방법은 $\widehat{C}(x) = \{y \in \mathcal{Y} \mid s(x,y) \leq q\}$ 형태의 예측 집합을 구성하며, $q$는 훈련 분포 주변의 $f$-divergence 구 내 모든 분포에 대해 $1-\alpha$ 커버리지를 보장하도록 선택된다.
  • 분포 이탈 하에서도 훈련 데이터의 교환 가능성 하에서 유한 표본 유효성을 확보하기 위해 분할 콕포멀 추론을 활용한다.
  • 검증 데이터로부터 유도된 일致한 추정기로 $f$-divergence 구의 크기를 추정함으로써 알려지지 않은 이격에 대한 강건성을 확보한다.
  • 경험적 모멘트 또는 공변수의 투영을 사용하여 불확실성 집합을 정의하는 히우리스틱이지만 잘 정당화된 이격 수준 추정 절차를 도입한다.
  • 예측 집합을 구성하고 캘리브레이션된 신뢰도를 확보하기 위해 점수 함수 $s(x,y)$를 사용하며, 예를 들어 음의 로그우도를 사용한다.
  • 유한 표본 조건 하에서도 $f$-divergence 구 내 모든 $Q$에 대해 거의 정확한 커버리지를 달성하며, 조건부 정규성 하에서 이격 추정기의 일致성도 증명한다.
Figure 3: Empirical coverage and average size for the prediction sets generated by the standard conformal methodology (“SC”) and the chi-squared divergence, across 20 random splits of the CIFAR-10 data. We set $\rho$ according to the sampling (“ $\chi^{2}$ -S”), regression (“ $\chi^{2}$ -R”), and cl
Figure 3: Empirical coverage and average size for the prediction sets generated by the standard conformal methodology (“SC”) and the chi-squared divergence, across 20 random splits of the CIFAR-10 data. We set $\rho$ according to the sampling (“ $\chi^{2}$ -S”), regression (“ $\chi^{2}$ -R”), and cl

실험 결과

연구 질문

  • RQ1훈련 분포 주변의 $f$-divergence 구 내에서 임의의 분포 이탈이 발생하더라도 유효한 커버리지를 유지하는 예측 집합을 구성할 수 있는가?
  • RQ2실제로 테스트 분포에 대한 사전 지식 없이도 분포 이격의 크기를 어떻게 추정하여 강건성을 캘리브레이션할 수 있는가?
  • RQ3분포 이탈 하에서 콕포멀 추론의 유한 표본 유효성은 어떻게 되며, 기존의 콕포멀 추론과 비교해 볼 때 어떤가?
  • RQ4제안된 방법은 CIFAR-v4와 ImageNet-V2와 같이 알려진 분포 이탈이 존재하는 실제 벤치마크에서 높은 신뢰도와 정확도를 유지할 수 있는가?
  • RQ5이 방법의 커버리지는 이격 추정 절차의 선택에 얼마나 민감한가? 그리고 강건성의 한계는 무엇인가?

주요 결과

  • 제안된 방법은 훈련 분포 주변의 $f$-divergence 구 내 모든 분포에 대해 (거의 정확한) $1-\alpha$ 커버리지를 유한 표본 조건 하에서도 달성한다.
  • 이 방법은 훈련 데이터의 교환 가능성만을 가정함으로써, i.i.d. 또는 파라미터 가정 없이도 유효성을 유지한다.
  • CIFAR-v4와 ImageNet-V2에서의 실증 결과는 테스트 데이터가 크게 이격된 경우에도 높은 커버리지를 유지하며, 분포 이탈 하에서 표준 콕포멀 방법보다 뛰어난 성능을 보인다.
  • 미약한 정규성 조건 하에서 이격 추정기의 일치성이 입증되어 실무에서 불확실성 집합의 신뢰성 있는 캘리브레이션을 가능하게 한다.
  • 민감도 분석 결과, 공변수 이격 하에서도 잘못된 커버리지가 유한하게 유지되며, 최악의 경우의 잘못된 커버리지는 추정된 $f$-divergence 수준에 의해 제어된다.
  • 실제 벤치마크 데이터셋에서 관찰된 자연 분포 이탈에 대해 강건성을 보이며, 표준 벤치마크에서 11% 이상의 정확도 하락이 발생하더라도 커버리지가 유지된다.
Figure 4: Empirical coverage and average size for the prediction sets generated by the standard conformal methodology (“SC”) and the chi-squared divergence, across 20 random splits of the MNIST data. We set $\rho$ according to the sampling (“ $\chi^{2}$ -S”), regression (“ $\chi^{2}$ -R”), and class
Figure 4: Empirical coverage and average size for the prediction sets generated by the standard conformal methodology (“SC”) and the chi-squared divergence, across 20 random splits of the MNIST data. We set $\rho$ according to the sampling (“ $\chi^{2}$ -S”), regression (“ $\chi^{2}$ -R”), and class

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.