[Paper Review] Large language models predict human sensory judgments across six modalities
State-of-the-art LLMs (GPT-3/3.5/4) produce pairwise sensory similarity judgments across six modalities that significantly correlate with human data, recovering known representations like the color wheel and pitch spiral, and revealing language-dependent effects in color naming.
Determining the extent to which the perceptual world can be recovered from language is a longstanding problem in philosophy and cognitive science. We show that state-of-the-art large language models can unlock new insights into this problem by providing a lower bound on the amount of perceptual information that can be extracted from language. Specifically, we elicit pairwise similarity judgments from GPT models across six psychophysical datasets. We show that the judgments are significantly correlated with human data across all domains, recovering well-known representations like the color wheel and pitch spiral. Surprisingly, we find that a model (GPT-4) co-trained on vision and language does not necessarily lead to improvements specific to the visual modality. To study the influence of specific languages on perception, we also apply the models to a multilingual color-naming task. We find that GPT-4 replicates cross-linguistic variation in English and Russian illuminating the interaction of language and perception.
Motivation & Objective
- Investigate how much perceptual information about the world can be recovered from language using large language models.
- Assess whether LLM-derived similarity judgments align with human perceptual representations across multiple modalities.
- Examine if multimodal training (text + images) or language alone drives modality-specific predictive power.
- Explore cross-linguistic effects in perception by testing color naming in English and Russian with LLMs.
Proposed method
- Elicit 10 pairwise similarity ratings per stimulus pair from GPT-3, GPT-3.5, and GPT-4 using tailored prompts and in-context examples.
- Compare model-derived similarity scores to human data using Pearson correlations across six modalities.
- Analyze the emergence of known perceptual structures via MDS to recover color wheel, pitch spiral, and consonant representations.
- Conduct a multilingual color-naming task (English and Russian) to test language dependence of perceptual representations.
- Provide model-generated explanations for judgments to assess alignment with perceptual concepts (octave relations, places of articulation, color spectra).

Experimental results
Research questions
- RQ1Can LLMs yield similarity judgments that align with human perceptual representations across multiple modalities?
- RQ2Do LLMs recover well-known perceptual structures such as the color wheel and pitch spiral from language?
- RQ3Does multimodal training improve modality-specific performance beyond language alone?
- RQ4Are color naming and perceptual representations influenced by prompt language, revealing language-dependent perception?
- RQ5To what extent do LLMs replicate cross-linguistic variation in color naming observed in humans?
Key findings
- GPT-4 shows strongest alignment with human data across most modalities, with correlations such as pitch r=.92 and colors r=.89.
- GPT-3.5 attains high correlations for loudness (r=.89) and other domains, with overall performance often in the top two models.
- Inter-rater reliability (IRR) for pitch (r=.90) and consonants (r=.46) suggests GPT-4 performance approaches human reliability in some domains.
- MDS analyses reveal interpretable perceptual spaces: a pitch spiral with 12-semitone structure, a color wheel, and production-based consonant representations.
- Color naming with GPT-4 replicates cross-linguistic differences between English and Russian, aligning with known human cross-language patterns.
- GPT-4’s enhanced performance is attributed to richer textual training rather than solely multimodal (image) inputs.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.