[논문 리뷰] ProPublica's COMPAS Data Revisited
이 논문은 ProPublica가 널리 사용하는 COMPAS 재범 데이터셋에서 중요한 데이터 처리 오류를 특정한다: ProPublica는 재범자에게 두 번째 연도 스크리닝 일자 마감일을 적용하지 않아, 두 번째 연도 재범률이 24.3% 증가했다(36.2%에서 45.1%로). 저자는 이 오류를 수정하기 위해 일관된 마감일을 적용하여, 핵심 공정성 지표인 PPV와 NPV가 크게 영향을 받는 것을 입증한다. 반면 FPR, FNR 및 정확도는 거의 변화가 없다.
I examine the COMPAS recidivism risk score and criminal history data collected by ProPublica in 2016 that fueled intense debate and research in the nascent field of 'algorithmic fairness'. ProPublica's COMPAS data is used in an increasing number of studies to test various definitions of algorithmic fairness. This paper takes a closer look at the actual datasets put together by ProPublica. In particular, the sub-datasets built to study the likelihood of recidivism within two years of a defendant's original COMPAS survey screening date. I take a new yet simple approach to visualize these data, by analyzing the distribution of defendants across COMPAS screening dates. I find that ProPublica made an important data processing error when it created these datasets, failing to implement a two-year sample cutoff rule for recidivists in such datasets (whereas it implemented a two-year sample cutoff rule for non-recidivists). When I implement a simple two-year COMPAS screen date cutoff rule for recidivists, I estimate that in the two-year general recidivism dataset ProPublica kept over 40% more recidivists than it should have. This fundamental problem in dataset construction affects some statistics more than others. It obviously has a substantial impact on the recidivism rate; artificially inflating it. For the two-year general recidivism dataset created by ProPublica, the two-year recidivism rate is 45.1%, whereas, with the simple COMPAS screen date cutoff correction I implement, it is 36.2%. Thus, the two-year recidivism rate in ProPublica's dataset is inflated by over 24%. This also affects the positive and negative predictive values. On the other hand, this data processing error has little impact on some of the other key statistical measures, which are less susceptible to changes in the relative share of recidivists, such as the false positive and false negative rates, and the overall accuracy.
연구 동기 및 목표
- 알고리즘 공정성 연구의 기초가 되는 널리 사용되는 ProPublica의 COMPAS 재범 데이터셋의 무결성을 조사하기 위해.
- 두 번째 연도 재범 데이터셋 제작 과정에서 발생한 심각한 데이터 처리 오류를 특정하고 수정하기 위해.
- 이 오류가 재범률, PPV, NPV, FPR/FNR와 같은 핵심 공정성 지표에 미치는 영향을 평가하기 위해.
- 알고리즘 공정성 연구를 위한 기준 데이터셋에서 데이터 품질과 처리 투명성의 중요성을 강조하기 위해.
- 수정된 데이터셋을 제공하고, 데이터셋의 제작 과정에 철저한 검토가 이루어지지 않은 채 사용할 경우의 위험을 경고하기 위해.
제안 방법
- 저자는 피고인들이 COMPAS 스크리닝 일자에 어떻게 분포되어 있는지 분석하여 데이터 처리의 일관성 문제를 탐지한다.
- 재범자와 비재범자 모두에 동일한 두 번째 연도 스크리닝 일자 마감일 규칙을 적용함으로써 ProPublica의 원래 비대칭적 처리 방식을 수정한다.
- 수정된 데이터셋을 사용하여 재범률, PPV, NPV, FPR, FNR, 정확도와 같은 핵심 통계치를 재계산한다.
- 원래 ProPublica의 통계와 수정된 데이터셋에서 유도된 통계를 비교하여 수정의 영향을 정량화한다.
- 분석은 알고리즘 공정성 연구에서 가장 자주 사용되는 두 번째 연도 일반 재범 데이터셋에 집중한다.
- 저자는 시각적 및 통계적 분석을 통해 원래 데이터셋이 재범자 비율이 과도하게 높은 이유가 마감일 규칙이 재범자에게 적용되지 않았기 때문임을 입증한다.
실험 결과
연구 질문
- RQ1ProPublica는 두 번째 연도 재범 데이터셋 제작 시 재범자와 비재범자 모두에게 일관된 두 번째 연도 스크리닝 일자 마감일 규칙을 적용했는가?
- RQ2결측한 마감일 규칙이 두 번째 연도 재범률에 미치는 정량적 영향은 무엇인가?
- RQ3이 데이터 처리 오류는 양성 예측도(PPV)와 음성 예측도(NPV)와 같은 핵심 공정성 지표에 어떻게 영향을 미치는가?
- RQ4거짓 양성률(FPR), 거짓 음성률(FNR), 정확도는 이 데이터 오류로 인해 어느 정도 영향을 받는가?
- RQ5이 데이터 처리 오류는 이 데이터셋을 사용한 알고리즘 공정성 연구의 타당성에 대해 어떤 광범위한 함의를 지닌다?
주요 결과
- ProPublica는 재범자에게 두 번째 연도 스크리닝 일자 마감일을 적용하지 않았지만, 비재범자에게는 적용하여, 두 번째 연도 데이터셋에서 재범자가 체계적으로 과대표현되는 결과를 초래했다.
- ProPublica의 원본 데이터셋에서 두 번째 연도 재범률은 8.8个百分点 증가하여 36.2%에서 45.1%로 상승했다.
- 이는 데이터 처리 오류로 인해 두 번째 연도 재범률이 상대적으로 24.3% 증가한 것이다.
- 양성 예측도(PPV)와 음성 예측도(NPV)는 오류의 영향을 크게 받는다. 이는 이 지표들이 재범자의 비율에 의존하기 때문이다.
- 거짓 양성률(FPR), 거짓 음성률(FNR), 전체 정확도는 거의 영향을 받지 않으며, 이는 이 지표들이 재범자 비율의 변화에 덜 민감하기 때문이다.
- 저자는 재범자에게 마감일 규칙이 적용되지 않아 ProPublica가 두 번째 연도 재범자를 44.3% 더 많이 유지했다고 추정한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.