[Paper Review] Assessing the Contribution of Semantic Congruency to Multisensory Integration and Conflict Resolution
This study investigates how semantic congruency—specifically gender-based consistency between audio and visual stimuli—affects audio-visual (AV) spatial localization and conflict resolution in complex, immersive environments. Using a virtual avatar setup with male and female avatars and voices, the authors demonstrate that while semantic congruency modulates the magnitude of the ventriloquism effect (visual bias toward semantically aligned stimuli), environmental statistics such as lip movement and motion remain the dominant factor in resolving AV conflicts.
The efficient integration of multisensory observations is a key property of the brain that yields the robust interaction with the environment. However, artificial multisensory perception remains an open issue especially in situations of sensory uncertainty and conflicts. In this work, we extend previous studies on audio-visual (AV) conflict resolution in complex environments. In particular, we focus on quantitatively assessing the contribution of semantic congruency during an AV spatial localization task. In addition to conflicts in the spatial domain (i.e. spatially misaligned stimuli), we consider gender-specific conflicts with male and female avatars. Our results suggest that while semantically related stimuli affect the magnitude of the visual bias (perceptually shifting the location of the sound towards a semantically congruent visual cue), humans still strongly rely on environmental statistics to solve AV conflicts. Together with previously reported results, this work contributes to a better understanding of how multisensory integration and conflict resolution can be modelled in artificial agents and robots operating in real-world environments.
Motivation & Objective
- To investigate the role of semantic congruency—specifically gender-based consistency—between audio and visual stimuli in multisensory integration and conflict resolution.
- To examine how semantic congruency influences the magnitude of the ventriloquism effect (perceptual shift of sound toward visual cues) in complex, real-world-like environments.
- To determine whether semantic factors can override or modulate the influence of spatial and dynamic visual statistics in AV conflict resolution.
- To provide empirical data for improving computational models of multisensory perception in artificial agents and robots operating in real-world settings.
Proposed method
- Conducted a behavioral study with 32 participants in an immersive projection environment featuring four animated avatars (three male, one female) producing spatially misaligned audio-visual stimuli.
- Presented audio stimuli with male or female voices and visual stimuli with corresponding or conflicting gender cues, while varying spatial alignment (central, lateral, 1- and 2-avatar gaps).
- Used a chin-rest to stabilize participant gaze and ensure consistent viewing distance (160 cm) from the screen, with loudspeakers positioned to match avatar locations.
- Measured participants’ responses to localize the source of auditory stimuli across conditions of spatial and semantic congruency.
- Analyzed data using repeated-measures ANOVAs to assess main effects and interactions of voice gender, visual cue type (static or animated), and spatial distance.
- Quantified the ventriloquism effect (ER) as the perceptual shift of sound toward visual cues, comparing conditions with and without semantic congruency.
Experimental results
Research questions
- RQ1How does semantic congruency (gender-specific audio-visual pairing) affect the magnitude of the ventriloquism effect in spatial audio-visual localization tasks?
- RQ2To what extent do environmental statistics (e.g., animated lip movements, body motion) influence AV conflict resolution compared to semantic congruency?
- RQ3Does the presence of a semantically congruent female avatar (FA) significantly bias perception when the sound is spatially aligned with a male avatar (MA)?
- RQ4How do spatial distances (1- and 2-avatar gaps) modulate the influence of semantic and dynamic visual cues on AV integration?
Key findings
- The main effect of voice gender on the ventriloquism effect was significant (F=7.82, p<0.01, η²=0.20), indicating that gender-specific audio-visual pairing influences perceptual bias.
- For male-voiced stimuli, the visual bias was significantly stronger in central (ER=70%) and lateral (ER=79%) conditions than in 1-avatar gap (ER=49%) and 2-avatar gap (ER=41%) conditions (p<0.001).
- For female-voiced stimuli, the animated male avatar (MA) induced a strong visual bias (ER=57%), while a static female avatar (FA) induced a much weaker bias (ER=10%), with both significantly different from animated FAs (ER=61%, p<0.001).
- The interaction between visual cue type and spatial distance was significant (F=5.61, p<0.01, η²=0.16), showing that animated avatars consistently induced stronger biases than static ones across distances.
- In conditions with male voices, the animated FA induced a higher bias (ER=52%) than the static FA (ER=56%) in the 1-avatar gap (p<0.01), suggesting that animation enhances perceptual binding even with semantic mismatch.
- Despite semantic congruency effects, environmental statistics (e.g., animated lip movements) were the dominant factor in AV conflict resolution, as evidenced by stronger biases toward animated avatars regardless of gender alignment.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.