[論文レビュー] HyFI: Hyperbolic Feature Interpolation for Brain-Vision Alignment
HyFIは双曲線的特徴補間を導入して意味的特徴と知覚的特徴を融合させ、モダリティギャップと絡み合いを解決し、THINGS-EEGおよび THINGS-MEG でのゼロショット脳-画像検索において最先端を達成します。
Recent progress in artificial intelligence has encouraged numerous attempts to understand and decode human visual system from brain signals. These prior works typically align neural activity independently with semantic and perceptual features extracted from images using pre-trained vision models. However, they fail to account for two key challenges: (1) the modality gap arising from the natural difference in the information level of representation between brain signals and images, and (2) the fact that semantic and perceptual features are highly entangled within neural activity. To address these issues, we utilize hyperbolic space, which is well-suited for considering differences in the amount of information and has the geometric property that geodesics between two points naturally bend toward the origin, where the representational capacity is lower. Leveraging these properties, we propose a novel framework, Hyperbolic Feature Interpolation (HyFI), which interpolates between semantic and perceptual visual features along hyperbolic geodesics. This enables both the fusion and compression of perceptual and semantic information, effectively reflecting the limited expressiveness of brain signals and the entangled nature of these features. As a result, it facilitates better alignment between brain and visual features. We demonstrate that HyFI achieves state-of-the-art performance in zero-shot brain-to-image retrieval, outperforming prior methods with Top-1 accuracy improvements of up to +17.3% on THINGS-EEG and +9.1% on THINGS-MEG.
研究の動機と目的
- 視覚的脳デコードにおける脳信号と言語表現のモダリティ間ギャップを動機付け、解決すること。
- 意味的特徴と知覚的特徴がニューラル活動において絡み合っており、個別に処理するのではなく統合すべきであることを示すこと。
- 意味的特徴と知覚的特徴の間を補間するために双曲幾何を活用し、脳の整合性を高めること。
- 双曲補間が表現を圧縮し、脳情報容量が限られていることを反映することを示すこと。
- 本法の一般的適用性をさまざまな視覚エンコーダ・脳エンコーダに対して確立すること。
提案手法
- Lorentz(双曲面)モデルと指数写像を用いて意味的特徴と知覚的特徴を双曲空間に埋め込むこと。
- 意味的特徴と知覚的特徴の間を双曲幾何学の測地線に沿って動的に学習される補間係数で補間すること。
- 脳信号を同じ双曲空間に射影し、双曲対比学習を適用して補間された画像埋め込みと整合させること。
- 双曲補間が補間埋め込みを原点へと集中させ、情報を圧縮することを示すこと。
- 補間された視覚表現と脳埋め込みの整合を課す双曲対比損失で訓練すること。
実験結果
リサーチクエスチョン
- RQ1意味的特徴と知覚的特徴の双曲補間は、ユークリッドアプローチと比較して脳信号(EEG/MEG)と視覚表現の整合性を改善できるか。
- RQ2双曲幾何学の測地線に沿った補間は、ニューラル活動における意味的・知覚的情報の絡み合いをよりよく捉えるか。
- RQ3HyFIは THINGS-EEG および THINGS-MEG のベンチマークでゼロショットの脳-画像検索においてどのように性能を示すか。
- RQ4異なる視覚エンコーダと脳エンコーダの組み合わせはHyFIの有効性にどのような影響を与えるか。
主な発見
- HyFIは THINGS-EEG でトップ1 68.2%、トップ5 91.9%というゼロショット脳-画像検索の最先端を達成。
- HyFIは THINGS-MEG でトップ1 35.8%、トップ5 64.6%というゼロショット脳-画像検索の最先端を達成。
- アブレーション研究は、双曲空間と双曲補間の組み合わせがユークリッド空間(CLIP)やユークリッド空間での補間を上回ることを示した。
- 双曲補間は補間埋め込みを原点へ集中させ、圧縮と表現容量の低下を反映する。
- HyFIは視覚エンコーダと脳エンコーダの組み合わせを横断して一貫して性能を向上させ、広い適用性を示す。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。