[논문 리뷰] Is There a Trade-Off Between Fairness and Accuracy? A Perspective Using Mismatched Hypothesis Testing
본 논문은 잘못 일치된 가설 검정과 Chernoff 정보를 통해 공정성-정확도 간의 trade-off를 재구성하고, 이상적 분포에서 고유한 trade-off가 없음을 보여주며, 실제 상황에서 trade-off를 완화하기 위한 기준을 제시한다.
A trade-off between accuracy and fairness is almost taken as a given in the existing literature on fairness in machine learning. Yet, it is not preordained that accuracy should decrease with increased fairness. Novel to this work, we examine fair classification through the lens of mismatched hypothesis testing: trying to find a classifier that distinguishes between two ideal distributions when given two mismatched distributions that are biased. Using Chernoff information, a tool in information theory, we theoretically demonstrate that, contrary to popular belief, there always exist ideal distributions such that optimal fairness and accuracy (with respect to the ideal distributions) are achieved simultaneously: there is no trade-off. Moreover, the same classifier yields the lack of a trade-off with respect to ideal distributions while yielding a trade-off when accuracy is measured with respect to the given (possibly biased) dataset. To complement our main result, we formulate an optimization to find ideal distributions and derive fundamental limits to explain why a trade-off exists on the given biased dataset. We also derive conditions under which active data collection can alleviate the fairness-accuracy trade-off in the real world. Our results lead us to contend that it is problematic to measure accuracy with respect to data that reflects bias, and instead, we should be considering accuracy with respect to ideal, unbiased data.
연구 동기 및 목표
- 공정성-정확도 문제에 동기를 부여하고 실제 데이터에서 가정된 trade-off에 도전한다.
- 그룹 간 정확성과 공정성을 정량하기 위해 Chernoff 정보를 통한 구분성(separability)을 도입한다.
- 관찰 데이터로의 편향된 매핑이 겉으로 보이는 trade-off를 만들 수 있음을 보인다.
- 공정성 및 정확성이 정렬되는 이상적 분포를 제안하고 그 구성 방법을 제시한다.
- 능동적 데이터 수집이 trade-off를 감소시키거나 제거하는 조건을 도출한다.
제안 방법
- 보호 속성 Z를 갖는 구성 공간과 편향된 관찰 공간에서 이진 분류를 모델링한다.
- 가능도 비 탐지기(likelihood ratio detectors)와 Chernoff 지수를 사용하여 각 그룹의 오류 확률을 정량화한다.
- 무권한 그룹과 특권 그룹 간의 P0/P1과 Q0/Q1 사이의 Chernoff 정보를 구분성으로 정의한다.
실험 결과
연구 질문
- RQ1관찰 데이터로의 구성 공간과 편향된 매핑 사이에 실제 세계의 정확도-공정성 trade-off가 발생하는가?
- RQ2공정성과 정확성이 동시에 최대로 달성될 수 있는 이상적 분포가 존재하는가?
- RQ3특성 데이터 수집 조건에서 특성과 수집특징을 늘리면 구분성이 향상되고 trade-off가 감소하는가?
- RQ4이상적 데이터를 통해 공정성을 유지하면서 정확성을 개선하는 이상적 분포를 어떻게 구성할 수 있는가?
주요 결과
- Chernoff 정보가 각 그룹에 대해 정확도-공정성 trade-off를 정량하는 구분성의 지표로 사용된다.
- If C(P0,P1) < C(Q0,Q1), the Bayes optimal detectors are unfair on the observed data, and any fairness adjustment lowers accuracy for at least one group (Theorem 1).
- There exist ideal distributions for the unprivileged group such that the Bayes optimal detector is fair on the given data and optimal on the ideal data (Theorem 2).
- An optimization framework can yield ideal distributions that minimize divergence from the observed data while achieving fairness and matching the privileged group’s separability on the ideal data (Theorem 2; optimization (4)).
- Active data collection can alleviate the trade-off by increasing separability (Theorem 3).
- The work argues accuracy should be evaluated with respect to ideal, unbiased data rather than biased observed data.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.