[논문 리뷰] It's COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks
이 논문은 프리트라이얼 RAI 데이터세트, 특히 COMPAS가 편향되어 있으며 벤치마킹 공정성에 대해 맥락적으로 불일치하고, 실세계 CJ 결과는 알고리즘 공정성 너머의 사회기술적 요인에 의존한다고 주장합니다. RAIs를 사용할 때 학제간 표준과 규범적 인식을 옹호합니다.
Risk assessment instrument (RAI) datasets, particularly ProPublica's COMPAS dataset, are commonly used in algorithmic fairness papers due to benchmarking practices of comparing algorithms on datasets used in prior work. In many cases, this data is used as a benchmark to demonstrate good performance without accounting for the complexities of criminal justice (CJ) processes. However, we show that pretrial RAI datasets can contain numerous measurement biases and errors, and due to disparities in discretion and deployment, algorithmic fairness applied to RAI datasets is limited in making claims about real-world outcomes. These reasons make the datasets a poor fit for benchmarking under assumptions of ground truth and real-world impact. Furthermore, conventional practices of simply replicating previous data experiments may implicitly inherit or edify normative positions without explicitly interrogating value-laden assumptions. Without context of how interdisciplinary fields have engaged in CJ research and context of how RAIs operate upstream and downstream, algorithmic fairness practices are misaligned for meaningful contribution in the context of CJ, and would benefit from transparent engagement with normative considerations and values related to fairness, justice, and equality. These factors prompt questions about whether benchmarks for intrinsically socio-technical systems like the CJ system can exist in a beneficial and ethical way.
연구 동기 및 목표
- 프리트라이얼 RAI 데이터셋의 편향과 오류를 강조하고 이것이 벤치마킹 공정성에 미치는 영향을 설명합니다.
- 범죄 사법 맥락에서 알고리즘 공정성만으로는 실세계의 공정성을 보장할 수 없다는 점을 설명합니다.
- RAI를 평가할 때 범죄학, 심리학, 법학, 윤리학의 학제간 기준을 촉진합니다.
- CJ에서 RAI 데이터셋을 사용하는 연구를 위한 규범적 고려사항과 실용적 지침을 제공합니다.
제안 방법
- 결과(Y), 보호속성(A), 공변량(X), 분포 전반에 걸친 편향과 오류를 조사합니다.
- CJ 절차 및 재량이 공정성 벤치마크의 실세계 결과 적용 가능성을 어떻게 제한하는지 논의합니다.
- 학제간 방법론 표준을 표준 ML 벤치마킹 관행과 비교합니다.
- 맥락적 정당화를 바탕으로 COMPAS/RAI 데이터를 사용하는 권고사항과 모범 사례를 제시합니다.
- AI 공정성 관행과 CJ 연구 규범 간의 불일치를 비판적으로 분석합니다.
실험 결과
연구 질문
- RQ1프리트라이얼 RAI 데이터세트에 벤치마킹 타당성을 약화시키는 편향과 측정 오류가 포함되어 있습니까?
- RQ2CJ 절차와 인간 재량이 알고리즘 예측을 넘어 실세계의 공정성에 어떤 영향을 미칩니까?
- RQ3RAI를 연구할 때 CJ의 학문적 표준(범죄학, 심리학, 법률)이 ML 벤치마킹 규범과 어떻게 다른가요?
- RQ4공정성 연구에서 RAI 데이터셋의 사용을 안내할 규범적 고려사항은 무엇입니까?
- RQ5CJ 연구에서 COMPAS 및 RAI 데이터를 의미있게 활용하기 위한 최선의 관행은 무엇입니까?
주요 결과
- COMPAS를 포함한 RAI 데이터셋은 결과, 보호 속성, 공변량 전반에 걸친 측정 편향과 오류를 포함하여 벤치마킹을 복잡하게 만듭니다.
- 분포 편향, 선택 효과, CJ 과정에서의 하류 재량이 알고리즘 공정성의 결과를 실세계 결과로 일반화하는 것을 제한합니다.
- 이전 RAI 실험을 단순히 재현하는 것은 공정성 개념과 CJ 맥락의 명시적 고찰 없이 규범적 입장을 강화할 수 있습니다.
- 학제간 참여는 사회기술적 CJ 시스템의 벤치마크가 맥락, 가치 및 윤리적 고려를 필요로 한다는 것을 보여줍니다.
- 현행 ML 출판 관행 및 데이터 중심 벤치마킹은 CJ 연구 목표와 맞지 않으며 실세계 영향력을 오해할 수 있습니다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.