Skip to main content
QUICK REVIEW

[Paper Review] Evaluating the Social Impact of Generative AI Systems in Systems and Society

Irene Solaiman, Zeerak Talat|arXiv (Cornell University)|Jun 9, 2023
Ethics and Social Impacts of AISocial Sciences43 citations
TL;DR

Proposes a framework to evaluate social impact of generative AI across modalities, separating base technical system evaluation from societal impact and outlining categories and methods.

ABSTRACT

Generative AI systems across modalities, ranging from text (including code), image, audio, and video, have broad social impacts, but there is no official standard for means of evaluating those impacts or for which impacts should be evaluated. In this paper, we present a guide that moves toward a standard approach in evaluating a base generative AI system for any modality in two overarching categories: what can be evaluated in a base system independent of context and what can be evaluated in a societal context. Importantly, this refers to base systems that have no predetermined application or deployment context, including a model itself, as well as system components, such as training data. Our framework for a base system defines seven categories of social impact: bias, stereotypes, and representational harms; cultural values and sensitive content; disparate performance; privacy and data protection; financial costs; environmental costs; and data and content moderation labor costs. Suggested methods for evaluation apply to listed generative modalities and analyses of the limitations of existing evaluations serve as a starting point for necessary investment in future evaluations. We offer five overarching categories for what can be evaluated in a broader societal context, each with its own subcategories: trustworthiness and autonomy; inequality, marginalization, and violence; concentration of authority; labor and creativity; and ecosystem and environment. Each subcategory includes recommendations for mitigating harm.

Motivation & Objective

  • Define social impact in the context of generative AI and motivate the need for standard evaluation across modalities.
  • Develop a two-part framework separating base system evaluations from people-and-society evaluations.
  • Identify and describe categories of social impact applicable to base systems and to society.
  • Propose methodologies and considerations for conducting these evaluations and mitigating harms.

Proposed method

  • Establish seven base-system categories of social impact (bias/stereotypes/representational harms; cultural values and sensitive content; disparate performance; privacy and data protection; financial costs; environmental costs; data and content moderation labor).
  • Define five society-focused overarching categories (trustworthiness/autonomy; inequality/marginalization/violence; concentration of authority; labor/creativity; ecosystem/environment) with subcategories and mitigation recommendations.
  • Present qualitative and quantitative evaluation approaches adaptable across modalities (text, image, video, audio) and highlight limitations of existing evaluations.
  • Describe a two-workshop methodology with expert input to build the framework and identify evaluation methods, plus ongoing CRAFT session for updates.
  • Discuss data, privacy, regulatory, and governance considerations affecting evaluation and the need for an open evaluation repository.

Experimental results

Research questions

  • RQ1What social impact categories are most relevant for base generative AI systems across modalities?
  • RQ2What societal-level impacts (trust, inequality, labor, governance) should be evaluated, and how can harms be mitigated?
  • RQ3How can evaluation methods be standardized, documented, and extended to future modalities and deployments?

Key findings

  • Proposes a structured two-part evaluation framework: technical base-system evaluations and people/society evaluations.
  • Catalogues seven base-system impact areas and five society-focused categories with subcategories and mitigation guidance.
  • Argues that evaluations should be both quantitative and qualitative to capture nuance and context.
  • Recognizes the absence of a universal governing body for social impact and the need for an ongoing, community-contributed evaluation repository.
  • Plans for an updated version of the framework informed by ACM FAccT 2023 discussions (CRAFT session).

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.