Skip to main content
QUICK REVIEW

[논문 리뷰] No computation without representation: Avoiding data and algorithm biases through diversity

Caitlin Kuhlman, Latifa Jackson|arXiv (Cornell University)|2020. 02. 26.
Ethics and Social Impacts of AI참고 문헌 61인용 수 17
한 줄 요약

이 논문은 데이터 과학 분야에서의 표현 부족이 데이터셋과 알고리즘에 구조적 편향을 야기하므로, 윤리적 AI를 달성하기 위해서는 계산 공동체 자체의 다양화가 필수적이라고 주장한다. 소수자 중심 기관과의 전략적 교육, 멘토링 및 협업을 통해 소외된 목소리를 통합함으로써, 초기 단계에서 공정성을 내재화하고 알고리즘 편향을 줄이며 공정한 사회기술 시스템을 조성할 수 있다.

ABSTRACT

The emergence and growth of research on issues of ethics in AI, and in particular algorithmic fairness, has roots in an essential observation that structural inequalities in society are reflected in the data used to train predictive models and in the design of objective functions. While research aiming to mitigate these issues is inherently interdisciplinary, the design of unbiased algorithms and fair socio-technical systems are key desired outcomes which depend on practitioners from the fields of data science and computing. However, these computing fields broadly also suffer from the same under-representation issues that are found in the datasets we analyze. This disconnect affects the design of both the desired outcomes and metrics by which we measure success. If the ethical AI research community accepts this, we tacitly endorse the status quo and contradict the goals of non-discrimination and equity which work on algorithmic fairness, accountability, and transparency seeks to address. Therefore, we advocate in this work for diversifying computing as a core priority of the field and our efforts to achieve ethical AI practices. We draw connections between the lack of diversity within academic and professional computing fields and the type and breadth of the biases encountered in datasets, machine learning models, problem formulations, and interpretation of results. Examining the current fairness/ethics in AI literature, we highlight cases where this lack of diverse perspectives has been foundational to the inequity in treatment of underrepresented and protected group data. We also look to other professional communities, such as in law and health, where disparities have been reduced both in the educational diversity of trainees and among professional practices. We use these lessons to develop recommendations that provide concrete steps for the computing community to increase diversity.

연구 동기 및 목표

  • 계산 및 데이터 과학 공동체의 다양성 부족이라는 알고리즘 편향의 근본 원인을 해결하기 위해 노력한다.
  • AI 연구에서의 소외된 표현이 공정성, 책임성 및 투명성의 맹점으로 이어지는 방식을 부각한다.
  • 다양한 시각이 데이터 및 모델 설계에서의 구조적 불평등을 식별하고 완화하는 데 필수적임을 입증한다.
  • 멘토링, 교육 협력 및 공동체 중심 연구 협업과 같은 실천 가능한 전략을 제안하여 더 포용적인 AI 연구 생태계를 조성한다.
  • 소외된 집단을 윤리적 AI의 리더로 육성하는 데 초점을 맞춘 하향식, 공동체 중심의 다양성 접근을 옹호한다.

제안 방법

  • 알고리즘 공정성 분야의 기존 문헌을 분석하여, 다양한 시각이 부족했던 사례들이 어떻게 편향된 모델 설계로 이어졌는지 규명한다.
  • 계산 분야의 소외와 데이터셋 및 알고리즘의 체계적 편향 사이의 유사성을 도출한다.
  • 법률 및 헬스케어 분야에서 성공한 다양성 프로그램을 분석하여 계산 분야의 최선의 실천 방안을 도출한다.
  • 멘토링 및 친화력 워크숍(예: 호러드 대학교에서 열린 BPDM 워크숍)이 공동체 형성과 기술 역량 향상에 어떻게 기여하는지 강조한다.
  • 소수자 중심 기관(MSIs)과의 교육 협력을 통해 AI 연구 파이프라인 내 표현을 증가시킨다.
  • 모델가 실제 사회적 형평성 문제를 반영할 수 있도록 도메인 전문가와의 연구 협력을 주장한다.

실험 결과

연구 질문

  • RQ1계산 연구에서 특정 인구 집단의 소외가 지속적인 알고리즘 편향을 초래하는 방식은 무엇인가?
  • RQ2동질적인 연구 팀이 데이터 및 모델 설계에서의 구조적 불평등을 탐지하거나 해결하지 못하는 이유는 무엇인가?
  • RQ3멘토링 및 공동체 구축 프로그램이 AI 연구의 다양성과 공정성에 미치는 역할은 무엇인가?
  • RQ4소수자 중심 기관 및 지역 전문가와의 협력이 AI 시스템의 공정성과 관련성 향상에 어떻게 기여하는가?
  • RQ5지속 가능한 AI 및 윤리적 컴퓨팅에서 다양성을 달성하기 위해 교육 및 연구 문화에 필요한 체계적 변화는 무엇인가?

주요 결과

  • 계산 공동체의 다양성 부족은 동질적인 팀이 구조적 불평등을 인식하거나 고려하지 못함으로써 편향된 데이터셋과 알고리즘의 발생에 직접 기여한다.
  • 특정 보호 속성이 직접 사용되지 않더라도, 다양한 시각은 알고리즘 시스템 내 간접적 및 명시적 차별을 식별하고 완화하는 데 필수적이다.
  • 멘토링 및 친화력 워크숍(예: 호러드 대학교에서 열린 BPDM 워크숍)은 공동체 형성, 기술 역량 향상 및 데이터 과학 분야에서 소외된 집단의 표현 증가에 성공적으로 기여한다.
  • 소수자 중심 기관과의 교육 협력은 AI 및 데이터 과학 분야에서 더 포용적인 인재 파이프라인을 조성하는 데 기여할 수 있다.
  • 도메인 전문가와의 공동체 기반 연구 협력은 더 공정하고 맥락에 맞는 모델 개발을 이끈다.
  • AI 연구 공동체의 다양화를 위한 의도적인 노력을 기울이지 않는 한, 알고리즘 조작만으로는 공정성을 달성하는 데 부족하고 일시적인 결과에 그칠 것이다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.