Skip to main content
QUICK REVIEW

[논문 리뷰] Foundation Models and Fair Use

Peter Henderson, Xuechen Li|arXiv (Cornell University)|2023. 03. 28.
Digital Rights Management and Security인용 수 7
한 줄 요약

이 논문은 미국의 공정 이용 규정 하에 저작권 보호 자료를 기반으로 기초 모델을 훈련시키는 데서 발생하는 법적 및 윤리적 위험을 분석하며, 실험을 통해 이러한 모델이 보호된 저작물과 매우 유사한 콘텐츠를 생성할 수 있음을 입증한다. 기술적 완화 조치를 제안하고, 법과 기술의 상호 공진화를 통해 혁신을 유지하면서도 준수를 보장할 것을 촉구한다.

ABSTRACT

Existing foundation models are trained on copyrighted material. Deploying these models can pose both legal and ethical risks when data creators fail to receive appropriate attribution or compensation. In the United States and several other countries, copyrighted content may be used to build foundation models without incurring liability due to the fair use doctrine. However, there is a caveat: If the model produces output that is similar to copyrighted data, particularly in scenarios that affect the market of that data, fair use may no longer apply to the output of the model. In this work, we emphasize that fair use is not guaranteed, and additional work may be necessary to keep model development and deployment squarely in the realm of fair use. First, we survey the potential risks of developing and deploying foundation models based on copyrighted content. We review relevant U.S. case law, drawing parallels to existing and potential applications for generating text, source code, and visual art. Experiments confirm that popular foundation models can generate content considerably similar to copyrighted material. Second, we discuss technical mitigations that can help foundation models stay in line with fair use. We argue that more research is needed to align mitigation strategies with the current state of the law. Lastly, we suggest that the law and technical mitigations should co-evolve. For example, coupled with other policy mechanisms, the law could more explicitly consider safe harbors when strong technical tools are used to mitigate infringement harms. This co-evolution may help strike a balance between intellectual property and innovation, which speaks to the original goal of fair use. But we emphasize that the strategies we describe here are not a panacea and more work is needed to develop policies that address the potential harms of foundation models.

연구 동기 및 목표

  • 미국의 공정 이용 규정 하에 인터넷에서 수집한 저작권 보호 자료를 기반으로 훈련된 기초 모델의 배포와 관련된 법적 위험을 분석하기 위해.
  • 이러한 모델의 생성 출력이 원본 저작권 보호 저작물의 시장 가치를 침해할 수 있는지 평가하기 위해.
  • 기초 모델을 공정 이용 원칙과 일치시키는 데 도움이 되는 기술적 완화 전략을 식별하기 위해.
  • 지적 재산권과 인공지능 혁신을 균형 있게 유지하기 위해 법과 기술 간의 상호 공진화적 접근을 옹호하기 위해.
  • 기계 학습 연구자와 법적 전문가들을 위한 실천 가능한 연구 및 정책 가이던스를 제공하기 위해.

제안 방법

  • 특히 변형 이용과 시장 영향의 맥락에서 공정 이용과 관련된 미국의 사례 법령을 조사하였다.
  • 기초 모델의 출력과 저작권 보호 훈련 데이터 사이의 높은 유사성을 입증하는 실험을 수행하였다.
  • 텍스트 생성(GPT), 코드 합성(Codex), 이미지 생성(Stable Diffusion)을 포함한 실제 응용 사례를 분석하였다.
  • 데이터 필터링, 워터마킹, 프롬프트 엔지니어링과 같은 기존 기술적 완화 전략을 평가하였다.
  • 기술 도구를 법적 기준과 일치시키는 프레임워크를 제안하였으며, 강력한 완화 조치가 이행된 경우 안전항구를 제안하였다.
  • 기술적 보호 조치를 공정 이용 보호의 근거로 인정할 수 있는 정책 메커니즘을 주장하였다.

실험 결과

연구 질문

  • RQ1기초 모델이 저작권 보호 훈련 데이터와 상당히 유사한 출력을 얼마나 많이 생성하는가?
  • RQ2기초 모델의 출력이 시장 경쟁력 있는 결과를 낳을 경우, 공정 이용 규정이 여전히 적용되는 조건은 무엇인가?
  • RQ3현재의 기술적 완화 전략은 저작권 침해 위험을 줄이는 데 얼마나 효과적인가?
  • RQ4기술적 보호 조치가 법적으로 공정 이용 보호의 근거로 인정될 수 있는가?
  • RQ5법과 기술 개발이 어떻게 상호 공진화하여 책임감 있는 기초 모델 배포를 지원할 수 있는가?

주요 결과

  • 실험 결과, GPT-3나 Stable Diffusion와 같은 대표적인 기초 모델이 저작권 보호 훈련 데이터와 매우 유사한 출력을 생성할 수 있음을 확인하였다.
  • 원본 저작물의 시장 가치를 복제하거나 대체하는 생성 출력은 시장 피해를 초래함으로써 공정 이용 방어가 약화될 수 있다.
  • 현재의 기술적 완화 조치, 예를 들어 데이터 필터링과 워터마킹은 공정 이용 준수를 확보하기 위해 단독으로는 부족하다.
  • 특히 출력이 원작 저작물의 경제적 가치에 영향을 주는 경우, 생성 기반 기초 모델에 대해 공정 이용 규정이 보장되지 않는다.
  • 강력한 기술적 보호 조치가 법적으로 인정된다면 공정 이용으로의 길이 열릴 수 있으며, 이는 정책 혁신의 필요성을 시사한다.
  • 공정 이용이 적용된다고 해도, 창작 및 노동 시장에서 데이터 창작자들에게 발생하는 중대한 피해는 기술적 해결책만으로는 해결되지 않는다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.