Skip to main content
QUICK REVIEW

[논문 리뷰] Cross-Entropy Loss Functions: Theoretical Analysis and Applications

Anqi Mao, Mehryar Mohri|arXiv (Cornell University)|2023. 04. 14.
Adversarial Robustness in Machine Learning인용 수 213
한 줄 요약

이 논문은 비점근적 H-일관성 경계를 비점근적 비판-합 손실의 광범위한 가족에 대해 제시하고, 이를 매끄러운 적대적 변형으로 확장하여 강건성을 향상시킨다. 이론적 보장과 광범위한 실증 평가를 모두 제공한다.

ABSTRACT

Cross-entropy is a widely used loss function in applications. It coincides with the logistic loss applied to the outputs of a neural network, when the softmax is used. But, what guarantees can we rely on when using cross-entropy as a surrogate loss? We present a theoretical analysis of a broad family of loss functions, comp-sum losses, that includes cross-entropy (or logistic loss), generalized cross-entropy, the mean absolute error and other cross-entropy-like loss functions. We give the first $H$-consistency bounds for these loss functions. These are non-asymptotic guarantees that upper bound the zero-one loss estimation error in terms of the estimation error of a surrogate loss, for the specific hypothesis set $H$ used. We further show that our bounds are tight. These bounds depend on quantities called minimizability gaps. To make them more explicit, we give a specific analysis of these gaps for comp-sum losses. We also introduce a new family of loss functions, smooth adversarial comp-sum losses, that are derived from their comp-sum counterparts by adding in a related smooth term. We show that these loss functions are beneficial in the adversarial setting by proving that they admit $H$-consistency bounds. This leads to new adversarial robustness algorithms that consist of minimizing a regularized smooth adversarial comp-sum loss. While our main purpose is a theoretical analysis, we also present an extensive empirical analysis comparing comp-sum losses. We further report the results of a series of experiments demonstrating that our adversarial robustness algorithms outperform the current state-of-the-art, while also achieving a superior non-adversarial accuracy.

연구 동기 및 목표

  • 교차 엔트로피를 대체 손실로 사용할 때 비점근적이고 가설 집합 특정 보장을 제공한다.
  • 로지스틱 손실과 일반화된 교차 엔트로피를 포함하는 광범위한 comp-sum 손실 가족을 특징짓는다.
  • 부드러운 적대적 컴-합 손실을 도입하고 적대적 설정에서 H-일관성 경계를 확립한다.
  • 정규화된 부드러운 적대적 컴-합 손실을 최소화하여 적대적 강건성 알고리즘을 개발한다.
  • 표준 데이터셋에서 comp-sum 손실을 실증적으로 비교하고 강건성 및 비적대적 정확도를 평가한다.

제안 방법

  • 컴-합 손실을 점수 차이의 합에 Phi2를 적용한 합성 Phi1으로 정의하고, 이를 로지스틱, 일반화된 교차 엔트로피, 그리고 평균 절대 오차(mean absolute error)까지 포괄하도록 한다.
  • Phi^tau 계열을 도입하여 comp-sum 손실을 매개하고 그 성질(오목성, 단조성, 리프시츠 연속성)을 도출한다.
  • 대칭적이고 완전한 가설 집합에 대한 H-일관성 경계를 증명하고, 대리손실과 0-1 손실을 연결하는 Gamma_tau 변환을 제시한다.
  • 최소화 가능성 간극을 분석하여 명시적이고 촘촘한 경계를 얻고 이 간극을 통해 손실 함수를 비교한다.
  • 국소적인 rho-일관성 아래에서 매끄러운(smooth) 적대적 컴-합 손실을 정의하고 H-일관성 경계를 증명함으로써 적대적 강건성으로 확장한다.
  • CIFAR-10, CIFAR-100, SVHN에서 comp-sum 손실과 적대적 강건성 알고리즘을 비교하는 실증 분석을 수행한다.

실험 결과

연구 질문

  • RQ1교차 엔트로피를 대체 손실로 사용할 때 어떤 비점근적이고 가설 집합 특유의 보장을 개발할 수 있는가?
  • RQ2로지스틱 및 일반화된 교차 엔트로피를 포함하는 광범위한 comp-sum 손실 가족으로 H-일관성 경계가 어떻게 확장되는가?
  • RQ3이들 경계에서의 최소화 가능성 간극의 역할은 무엇이며 손실 간 비교는 어떻게 다른가?
  • RQ4매끄러운(smooth) 적대적 comp-sum 손실이 증명 가능한 H-일관성 경계와 향상된 적대적 강건성을 제공할 수 있는가?
  • RQ5표준 데이터셋에 대한 실증 결과가 comp-sum 손실과 그 적대적 변형의 이론적 이점을 지지하는가?

주요 결과

  • 다중 클래스 분류에서 로지스틱 손실에 대한 최초의 H-일관성 경계가 도출되었다.
  • 경계는 Gamma_tau 변환과 손실 및 가설 집합에 의존하는 최소화 가능성 간극을 통해 표현된다.
  • 구성적 논증을 통해 경계가 좁혀진다(타이트함이 보장된다).
  • 매끄러운_smooth 적대적 컴-합 손실은 H-일관성 경계를 허용하고 적대적 강건성 알고리즘을 가능하게 한다.
  • 실증 분석에서 매끄러운 적대적 comp-sum 손실에 기반한 적대적 알고리즘이 최첨단 기준치를 상회하고 비적대적 정확도도 향상시킨다.
  • tau 값에 따른 최소화 가능성 간극의 거동(로지스틱 및 평균 절대 오차 포함)을 특징화하고 이를 통해 손실 간 비교에 활용한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.