[논문 리뷰] Confidence-Building Measures for Artificial Intelligence: Workshop Proceedings
이 워크숍은 기초 모델로 인한 보안 위험을 완화하기 위한 실용적인 신뢰 구축 조치를 식별하며, 다중 이해관계자의 참여와 적응 가능한 비강제적 조치를 강조한다.
Foundation models could eventually introduce several pathways for undermining state security: accidents, inadvertent escalation, unintentional conflict, the proliferation of weapons, and the interference with human diplomacy are just a few on a long list. The Confidence-Building Measures for Artificial Intelligence workshop hosted by the Geopolitics Team at OpenAI and the Berkeley Risk and Security Lab at the University of California brought together a multistakeholder group to think through the tools and strategies to mitigate the potential risks introduced by foundation models to international security. Originating in the Cold War, confidence-building measures (CBMs) are actions that reduce hostility, prevent conflict escalation, and improve trust between parties. The flexibility of CBMs make them a key instrument for navigating the rapid changes in the foundation model landscape. Participants identified the following CBMs that directly apply to foundation models and which are further explained in this conference proceedings: 1. crisis hotlines 2. incident sharing 3. model, transparency, and system cards 4. content provenance and watermarks 5. collaborative red teaming and table-top exercises and 6. dataset and evaluation sharing. Because most foundation model developers are non-government entities, many CBMs will need to involve a wider stakeholder community. These measures can be implemented either by AI labs or by relevant government actors.
연구 동기 및 목표
- 오해와 고조를 방지하기 위해 기초 모델 시대에 신뢰 구축 조치(CBMs)의 필요성을 고취한다.
- 실행 가능한 CBMs의 집합을 식별하여 연구소, 정부, 시민사회를 비롯한 다양한 행위자에 적용 가능한 AI 시스템에 적용 가능하도록 한다.
- CBMs가 급속한 AI 혁신을 관리하기 위해 형식적 규제와 함께 어떻게 작동할 수 있는지 설명한다.
- CBM의 성공과 채택에 영향을 미칠 수 있는 정치적·기술적 한계를 강조한다.
- 기존 AI 거버넌스 프레임워크에 CBMs를 통합하기 위한 경로를 제안한다.
제안 방법
- 기초 모델에 적용되는 CBMs를 식별한다. 위기 핫라인, 사고 공유, 모델/시스템 카드를 포함하고, 콘텐츠 출처 확인 및 워터마크, 협력적 레드팀핑, 탑다운 시나리오 훈련, 데이터/평가 공유를 포함한다.
- CBMs를 의사소통 및 조정, 관찰 및 검증, 협력 및 통합, 투명성의 네 가지 범주로 정리한다.
- CBMs 구현에서 비정부 주체와 다중 이해관계자 참여의 역할을 논의한다.
- 역사적 및 현대 국제 보안 맥락의 사례와 고려사항을 제시한다.
- 한계와 지속적인 레드팀 수행 및 거버넌스 정렬의 필요성을 평가한다.
실험 결과
연구 질문
- RQ1기초 모델에서 국제 보안 위험을 완화하는 데 가장 적용 가능한 CBMs는 무엇인가?
- RQ2많은 AI 개발자가 비정부 주체이고 다중 이해관계자 참여가 필요하다는 점을 고려할 때 CBMs는 어떻게 구현될 수 있는가?
- RQ3AI를 위한 CBMs의 실행 가능성과 효과에 영향을 미치는 정치적·기술적 한계는 무엇인가?
- RQ4CBMs가 기존의 국제 규제 논의와 프레임워크를 어떻게 보완할 수 있는가?
주요 결과
- 기초 모델에 적용 가능한 CBMs로 위기 핫라인, 사고 공유, 모델/투명성/시스템 카드, 콘텐츠 출처 확인 및 워터마크, 협력적 레드팀핑, 탁상 연습, 데이터셋/평가 공유가 식별된다.
- CBMs는 의사소통/조정, 관찰/검증, 협력/통합, 그리고 투명성으로 분류되며 오해와 고조를 줄이도록 설계된다.
- 많은 CBMs는 자발적이며 AI 연구소나 정부 주체에 의해 시행될 수 있으며, 다수의 개발자가 비정부 주체인 특성으로 인해 다중 이해관계자 참여의 가능성이 있다.
- CBMs에는 검증 도전, 인센티브 정렬, 그리고 적응 가능한 구축형 접근을 요구하는 AI 능력의 진화와 같은 정치적·기술적 한계가 있다.
- 제안된 CBMs는 형식적 규제 노력을 보완할 수는 있지만 대체할 수 없으며, 신뢰가 낮은 국제 환경에서 다리 역할을 할 수 있다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.