Skip to main content
QUICK REVIEW

[논문 리뷰] Large Language Models Reflect the Ideology of their Creators

Maarten Buyl, Alexander Rogiers|arXiv (Cornell University)|2024. 10. 24.
Computational and Text Analysis Methods인용 수 10
한 줄 요약

이 논문은 17개의 인기 LLM의 이데올로기 입장을 영어와 중국어로 분석하고, 언어 프롬프트와 제작자 지역이 모델 이념에 영향을 준다는 것을 보여주며, 모델 간 차이가 상당히 크다. 설계 선택에 대한 투명성을 옹호하고 중립성을 가정하는 것을 경계한다.

ABSTRACT

Large language models (LLMs) are trained on vast amounts of data to generate natural language, enabling them to perform tasks like text summarization and question answering. These models have become popular in artificial intelligence (AI) assistants like ChatGPT and already play an influential role in how humans access information. However, the behavior of LLMs varies depending on their design, training, and use. In this paper, we prompt a diverse panel of popular LLMs to describe a large number of prominent personalities with political relevance, in all six official languages of the United Nations. By identifying and analyzing moral assessments reflected in their responses, we find normative differences between LLMs from different geopolitical regions, as well as between the responses of the same LLM when prompted in different languages. Among only models in the United States, we find that popularly hypothesized disparities in political views are reflected in significant normative differences related to progressive values. Among Chinese models, we characterize a division between internationally- and domestically-focused models. Our results show that the ideological stance of an LLM appears to reflect the worldview of its creators. This poses the risk of political instrumentalization and raises concerns around technological and regulatory efforts with the stated aim of making LLMs ideologically 'unbiased'.

연구 동기 및 목표

  • LLMs가 언어와 지역을 가로질러 창작자의 이데올로기를 반영하는지 조사한다.
  • 다양한 LLM이 생성하는 논란의 여지가 있는 역사적 인물에 대한 도덕적 평가를 정량화한다.
  • 프롬프트 언어(영어 vs 중국어)가 LLM의 이념적 입장에 미치는 영향을 살펴본다.
  • 이념 측면에서 서구 모델과 비서구 모델 간의 차이를 평가한다.
  • 규제, 투명성 및 모델 개발에 대한 시사점을 논의한다.

제안 방법

  • 두 단계의 자유로운 개방형 탐출: 1단계에서 LLM이 정치인을 서술하고, 2단계에서 LLM이 1단계 텍스트에 나타난 도덕적 평가를 평가하도록 한다.
  • 영어와 중국어로 평가된 17개 LLM 패널(표 2에 나열).
  • Pantheon 데이터세트(4,339개 인물)에서 다중 기준 필터링 및 인기 임계값으로 정치인을 선택.
  • 정치적 성향 해석 용이성을 높이기 위해 Manifesto Project 태그(61개 태그)로 주석.
  • Stage 1 설명을 Wikipedia 요약과 연결하고 Stage 2가 Likert 척도 프롬프트를 따르는지 데이터 품질 점검.
Figure 2 : Biplot showing the two-dimensional PCA-projection of the respondent’s average score for each ideology tag, with the factor loadings visualized as a grey vector that has a thickness proportional to the loading’s norm. To clarify the effect of the prompting language, Chinese respondents are
Figure 2 : Biplot showing the two-dimensional PCA-projection of the respondent’s average score for each ideology tag, with the factor loadings visualized as a grey vector that has a thickness proportional to the loading’s norm. To clarify the effect of the prompting language, Chinese respondents are

실험 결과

연구 질문

  • RQ1정치인을 기술할 때 LLM이 언어(영어 대 중국어)에 따라 체계적인 이념 차이를 보이는가?
  • RQ2서구 LLM과 비서구 LLM이 정치인에 대한 평가 및 자유민주적 가치와의 정합성에서 차이가 있는가?
  • RQ3프롬프트 언어와 모델 기원이 특정 이데올로기 및 정치 행위자에 대한 태도 형성에 어떻게 상호작용하는가?
  • RQ4개방형 도출을 사용하여 광범위한 LLM 세트에서 이념적 다양성을 정량화하고 시각화할 수 있는가?

주요 결과

  • 중국어 프롬프트는 일반적으로 중국 친화 인물 및 중앙집권적 거버넌스 특성에 대해 더 호의적인 시각을 낳는다.
  • 서구 모델은 영어로 프롬프트될 때 비서구 모델보다 자유민주적 가치와 인권 관련 태그를 더 긍정적으로 평가하는 경향이 있다.
  • 서구 모델들 간에도 OpenAI, Gemini, Mistral, Anthropic 계열 사이에 이념적 차이가 뚜렷하며 포용성, 거버넌스, 부패에 대한 서로 다른 강조가 있다.
  • 프롬프트 언어가 LLM 이념성의 분산의 상당 부분을 설명한다(중국어 vs 영어 차이에 대해 p = 0.0008).
  • 비서구 모델은 중앙집권적 경제 거버넌스와 국가 안정을 비교적 더 지지하는 반면, 서구 모델은 개인의 자유와 사회 정의를 더 선호한다.
  • 모델 간에 언어 간(영어 대 중국어) 및 지역 간(서구 대 비서구) 이념적 정렬의 증거가 있다.
Figure 3 : Average score difference over all respondents prompted in Chinese versus English. Red line indicates overall mean difference. Only the top 20 most positive and top 20 most negative differences are shown.
Figure 3 : Average score difference over all respondents prompted in Chinese versus English. Red line indicates overall mean difference. Only the top 20 most positive and top 20 most negative differences are shown.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.