Skip to main content
QUICK REVIEW

[논문 리뷰] Evaluating Model Performance Under Worst-case Subpopulations

Mike Li, Mittal, Daksh|arXiv (Cornell University)|2024. 07. 01.
Bayesian Modeling and Causal Inference인용 수 6
한 줄 요약

이 논문은 특정 크기의 모든 부분집합으로 구성된 모든 부분집합에 대해 최악의 경우 서브인구(performance)를 정의하고 추정하며, 편향되지 않고 확장 가능한 유한 샘플 보장을 제공하는 평가 방법을 제시한다.

ABSTRACT

The performance of ML models degrades when the training population is different from that seen under operation. Towards assessing distributional robustness, we study the worst-case performance of a model over all subpopulations of a given size, defined with respect to core attributes Z. This notion of robustness can consider arbitrary (continuous) attributes Z, and automatically accounts for complex intersectionality in disadvantaged groups. We develop a scalable yet principled two-stage estimation procedure that can evaluate the robustness of state-of-the-art models. We prove that our procedure enjoys several finite-sample convergence guarantees, including dimension-free convergence. Instead of overly conservative notions based on Rademacher complexities, our evaluation error depends on the dimension of Z only through the out-of-sample error in estimating the performance conditional on Z. On real datasets, we demonstrate that our method certifies the robustness of a model and prevents deployment of unreliable models.

연구 동기 및 목표

  • distribution shifts across arbitrary Z-defined subpopulations에서 ML 모델의 견고한 평가를 동기화한다.
  • W_alpha* 및 배치 안전을 위한 증명(alpha*)를 정의한다.
  • 조건부 위험 및 꼬리 위험 목표를 근사하기 위한 확장 가능 2단계 추정 절차를 개발한다.
  • Z에 대해 차원 독립적인 유한 샘플 및 점근 보장을 제공하여 mu 추정에 딥 네트워크를 사용할 수 있게 한다.

제안 방법

  • worst-case subpopulation performance W_alpha*를 conditional risk mu*(Z)의 tail-average(CVaR)로 형식화한다.
  • 듀얼 재구성을 적용하여 W_alpha*를 스칼라 eta와 plus-term의 최소화로 표현하고 mu*(Z)의 (1-alpha)-분위수와 관련짓는다.
  • auxiliary data에서 모델 클래스 H에 대해 회귀형 문제를 해결하여 conditional risk mu*(Z)을 추정한다.
  • mu*(Z)을 최종 W_alpha* 계산과 관련해 1차 오차를 보정하는 debiased(augmented) 추정기를 사용한다.
  • 다수의 Fold를 결합하고 중앙극한정리를 얻기 위해 cross-fitting을 사용하여 로버스트하고 데이터 효율적인 추정기를 만든다.
  • robustness certificate alpha*와 그 신뢰구간을 추정하는 절차를 제공한다.

실험 결과

연구 질문

  • RQ1임의의 subpopulation 정의 Z에 대해 최악의 경우 서브인구 성능을 어떻게 양성화하고 인증할 수 있는가?
  • RQ2tail-risk 목표 W_alpha*에 대해 편향되지 않고 확장 가능한 추정기를 구성할 수 있으며 유한 샘플 보장을 얻을 수 있는가?
  • RQ3수렴 속도는 무엇이며 조건부 위험 모델 클래스 H의 복잡도에 어떻게 의존하는가?
  • RQ4퍼포먼스가 허용 가능한 최소 서브인구 크기를 나타내는 임계값 alpha*를 통해 강건성을 어떻게 인증할 수 있는가?

주요 결과

  • debias된 2단계 추정기가 worst-case subpopulation performance에 대해 O_p(sqrt(Comp_n(H)/n))의 수렴 속도를 달성한다.
  • 방법은 mu*(Z) 추정의 외부 샘플 오차에 따라 bound가 의존하는 차원 독립적 농도 경계(concentration bound)를 허용한다.
  • Central limit theorem은 mû이 느리게 수렴하더라도 debiased 추정기에 대해 sqrt(n) 속도를 보인다.
  • 이 접근법은 worst-case subpopulation performance를 conditional value-at-risk 및 일관된 위험 측정과 연결한다.
  • 이 방법론은 실용적 강건성 증명 alpha*를 지원하고 mu 추정을 위한 딥 네트워크를 사용한 모델 평가를 가능하게 한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.