[논문 리뷰] HyFI: Hyperbolic Feature Interpolation for Brain-Vision Alignment
HyFI는 의미적 특징과 지각적 특징을 하이퍼볼릭 공간에서 보간하여 모달리티 간 차이와 얽힘을 해결하고 THINGS-EEG 및 THINGS-MEG에서 제로샷 뇌-이미지 검색에서 최첨단 성능을 달성한다.
Recent progress in artificial intelligence has encouraged numerous attempts to understand and decode human visual system from brain signals. These prior works typically align neural activity independently with semantic and perceptual features extracted from images using pre-trained vision models. However, they fail to account for two key challenges: (1) the modality gap arising from the natural difference in the information level of representation between brain signals and images, and (2) the fact that semantic and perceptual features are highly entangled within neural activity. To address these issues, we utilize hyperbolic space, which is well-suited for considering differences in the amount of information and has the geometric property that geodesics between two points naturally bend toward the origin, where the representational capacity is lower. Leveraging these properties, we propose a novel framework, Hyperbolic Feature Interpolation (HyFI), which interpolates between semantic and perceptual visual features along hyperbolic geodesics. This enables both the fusion and compression of perceptual and semantic information, effectively reflecting the limited expressiveness of brain signals and the entangled nature of these features. As a result, it facilitates better alignment between brain and visual features. We demonstrate that HyFI achieves state-of-the-art performance in zero-shot brain-to-image retrieval, outperforming prior methods with Top-1 accuracy improvements of up to +17.3% on THINGS-EEG and +9.1% on THINGS-MEG.
연구 동기 및 목표
- 시각적 뇌 해독에서 뇌 신호와 이미지 표현 간의 모달리티 차이를 동기 부여하고 해결한다.
- 의미적 특징과 지각적 특징이 신경 활동에서 얽혀 있으며 서로 분리해 처리하기보다 융합되어야 함을 보인다.
- 뇌 정합성을 높이기 위해 의미적 특징과 지각적 특징 사이를 하이퍼볼릭 기하학으로 보간하는 것을 활용한다.
- 하이퍼볼릭 특징 보간이 표현을 압축하여 뇌의 정보 용량이 제한되었음을 반영함을 보여준다.
- 다양한 시각 인코더와 뇌 인코더에 걸친 방법의 일반적 적용 가능성을 확립한다.
제안 방법
- Lorentz(하이퍼볼로이드) 모형과 지수 맵을 사용하여 의미적 및 지각적 이미지 특징을 하이퍼볼릭 공간에 임베딩한다.
- 동적으로 학습된 보간 계수를 갖고 하이퍼볼릭 기하학적 지오데식(측지선) 따라 의미적 및 지각적 특징을 보간한다.
- 뇌 신호를 동일한 하이퍼볼릭 공간으로 투영하고 하이퍼볼릭 대조 학습을 적용하여 보간된 이미지 임베딩과 정렬시킨다.
- 하이퍼볼릭 보간이 표현을 원점 쪽으로 집중시켜 정보를 효과적으로 압축함을 보인다.
- 보간된 시각 표현과 뇌 임베딩 간의 정렬을 강제하는 하이퍼볼릭 대비 손실로 학습한다.
실험 결과
연구 질문
- RQ1의미적 및 지각적 특징의 하이퍼볼릭 보간이 뇌 신호(EEG/MEG)와 시각적 표현 간의 정렬을 유클리드 접근법과 비교하여 개선할 수 있는가?
- RQ2하이퍼볼릭 기하학적 지오데식을 따라 보간하는 것이 신경 활동에서 얽힌 의미적 및 지각적 정보의 특성을 더 잘 포착하는가?
- RQ3THINGS-EEG 및 THINGS-MEG 벤치마크에서 HyFI가 제로샷 뇌-이미지 검색에서 어떻게 수행하는가?
- RQ4다양한 시각 인코더와 뇌 인코더가 HyFI의 효과성에 미치는 영향은 무엇인가?
주요 결과
- HyFI는 THINGS-EEG에서 Top-1 68.2% 및 Top-5 91.9%로 제로샷 뇌-이미지 검색에서 최첨단을 달성한다.
- HyFI는 THINGS-MEG에서 Top-1 35.8% 및 Top-5 64.6%로 제로샷 뇌-이미지 검색에서 최첨단을 달성한다.
- 변형 실험은 하이퍼볼릭 공간과 하이퍼볼릭 보간이 유클리드(CLIP) 공간 및 유클리드 공간의 보간을 능가함을 보여준다.
- 하이퍼볼릭 보간은 보간된 임베딩을 원점 쪽으로 집중시켜 압축 및 표현 용량 감소를 반영한다.
- HyFI는 시각 및 뇌 인코더의 조합 전반에서 지속적으로 성능을 향상시키며 광범위한 적용 가능성을 보여준다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.