[논문 리뷰] Gromov-Wasserstein unsupervised alignment reveals structural correspondences between the color similarity structures of humans and large language models
이 논문은 Gromov-Wasserstein 최적 수송을 이용한 비감독 정렬을 통해 인간과 두 GPT 모델(GPT-3.5 및 GPT-4)의 색상 유사성 구조를 비교하고, GPT-4의 구조가 색상-신경형 인간과 미세 항목 수준에서 밀접하게 정렬된다는 것을 발견했다.
Large Language Models (LLMs), such as the General Pre-trained Transformer (GPT), have shown remarkable performance in various cognitive tasks. However, it remains unclear whether these models have the ability to accurately infer human perceptual representations. Previous research has addressed this question by quantifying correlations between similarity response patterns of humans and LLMs. Correlation provides a measure of similarity, but it relies pre-defined item labels and does not distinguish category- and item- level similarity, falling short of characterizing detailed structural correspondence between humans and LLMs. To assess their structural equivalence in more detail, we propose the use of an unsupervised alignment method based on Gromov-Wasserstein optimal transport (GWOT). GWOT allows for the comparison of similarity structures without relying on pre-defined label correspondences and can reveal fine-grained structural similarities and differences that may not be detected by simple correlation analysis. Using a large dataset of similarity judgments of 93 colors, we compared the color similarity structures of humans (color-neurotypical and color-atypical participants) and two GPT models (GPT-3.5 and GPT-4). Our results show that the similarity structure of color-neurotypical participants can be remarkably well aligned with that of GPT-4 and, to a lesser extent, to that of GPT-3.5. These results contribute to the methodological advancements of comparing LLMs with human perception, and highlight the potential of unsupervised alignment methods to reveal detailed structural correspondences. This work has been published in Scientific Reports, DOI: https://doi.org/10.1038/s41598-024-65604-1.
연구 동기 및 목표
- LLM이 상관 분석을 넘어서 인간 지각 색상 표현을 추론하는지 평가한다.
- 인간 색상 유사성 구조와 LLM 간의 미세 항목 수준의 구조적 정렬을 평가한다.
- GPT-4와 GPT-3.5를 색상 공간 기저(RGB, LAB) 및 색상 비정형 참여자와 비교한다.
- GWOT 기반의 비감독 정렬의 유용성을 보이며 표현적 등가성을 검토한다.
제안 방법
- 색상-신경형(n=426) 및 색상-비정형(n=207) 참가자들로부터 93색상에 대한 색상 유사성 판단을 수집한다.
- hex로 인코딩된 색상 프롬프트와 각 쌍당 5회의 시도를 사용하여 GPT-3.5(gpt-3.5-turbo)와 GPT-4(gpt-4-0314)로부터 색상 유사성 판단을 얻고 결과를 평균화한다.
- 유클리드 거리 및 delta_E_cie2000에 따라 RGB 및 LAB 색상 공간 유사도 행렬을 생성한다.
- 유사도 행렬 간의 기존 RSA 상관을 계산한다.
- 레이블 대응을 가정하지 않고 GWOT(그로모프-와서스타인 최적 수송)을 적용하여 유사성 구조를 정렬한다; 엔트로피 정규화(epsilon [1e-4, 1e-1])로 최적화한다.
- GW 수송 계획을 사용하여 상위 1 및 상위 k 매칭 비율로 정렬을 평가한다.

실험 결과
연구 질문
- RQ1비감독 GWOT 정렬이 인간 색상 유사성 구조와 LLM 간의 미세 항목 수준 대응을 전통적 상관관계를 넘어 밝혀낼 수 있는가?
- RQ2어떤 모델들이(GPT-4, GPT-3.5, RGB, LAB) 인간 색상 지각과 구조적으로 가장 밀접하게 정렬되는가?
- RQ3색상-비정형 참여자들은 LLM이나 색상 공간 모델과 색상-신경형 참가자와 유사하게 정렬되는가?
- RQ4비감독 정렬을 고려할 때 GPT-4의 색상 구조가 GPT-3.5보다 인간과 더 많이 정렬되는가?
주요 결과
- 색상-신경형 인간의 색상 구조가 비감독 GWOT 정렬 하에서 GPT-4와 매우 잘 정렬되며(상위 1위 약 80%), 인간-인간 정렬에 근접한다(약 86%).
- GPT-3.5는 GPT-4에 비해 정렬이 낮다(상위 1위 약 31%).
- 색상 공간 모델(RGB: 약 3.8% 상위 1위; LAB: 약 7.0% 상위 1위)은 인간의 색상 구조와의 비감독 정렬이 거의 보이지 않으며, RSA 상관은 비교적 중간 수준이다(RGB 0.60, LAB 0.71).
- GPT-4의 비감독 정렬은 GPT-3.5 및 색상 공간 기저를 능가하며, 상관관계만으로는 포착되지 않는 미묘한 구조적 유사성을 보여준다.
- 색상-비정형 참가자들은 상관이 낮고 정렬도 더 불량하여 신경형 인간 및 LLM과 다른 색상 유사성 구조를 보인다.
- 정렬된 임베딩의 시각화는 신경형 인간과 GPT-4 간에 비슷한 색상들이 서로 클러스터링되는 것을 보여준다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.