[論文レビュー] CauCLIP: Bridging the Sim-to-Real Gap in Surgical Video Understanding via Causality-Inspired Vision-Language Modeling
CauCLIP は、ターゲットドメインデータなしで sim-to-real のドメインシフトに一般化する頑健な手術フェーズ認識のための因果性ガイド付き CLIP ベースのフレームワークを提案。周波数ベースの augmentation と因果抑制損失を用いる。
Surgical phase recognition is a critical component for context-aware decision support in intelligent operating rooms, yet training robust models is hindered by limited annotated clinical videos and large domain gaps between synthetic and real surgical data. To address this, we propose CauCLIP, a causality-inspired vision-language framework that leverages CLIP to learn domain-invariant representations for surgical phase recognition without access to target domain data. Our approach integrates a frequency-based augmentation strategy to perturb domain-specific attributes while preserving semantic structures, and a causal suppression loss that mitigates non-causal biases and reinforces causal surgical features. These components are combined in a unified training framework that enables the model to focus on stable causal factors underlying surgical workflows. Experiments on the SurgVisDom hard adaptation benchmark demonstrate that our method substantially outperforms all competing approaches, highlighting the effectiveness of causality-guided vision-language models for domain-generalizable surgical video understanding.
研究の動機と目的
- 限られた実データのアノテーションに起因する手術フェーズ認識の sim-to-real ドメインギャップに対処する。
- CLIP を基盤とした因果性を取り入れたビジョン-言語フレームワークを開発し、ドメイン非依存な表現を学習する。
- 周波数ベースの augmentation と因果抑制損失を導入して非因果的バイアスを低減する。
提案手法
- 手術フェーズ認識をビデオ-テキスト整合性を用いたビデオ-テキストマッチングタスクとして CLIP ベースで構築する。
- 周波数領域の augmentation を導入し、高周波成分の非因果的手掛かりを撹乱しつつ意味を保持する。
- 元の表征と周波数拡張表現の類似性を強制し、非因果的特徴を相互相関除去する因果抑制モジュールを実装する。
- 拡張ビュー間の意味的一貫性を維持するための拡張付き整合性損失を追加する。
- 元の CLIP 損失、拡張整合損失、抑制損失を組み合わせた総損失で学習する。
実験結果
リサーチクエスチョン
- RQ1ターゲットドメインアクセスなしで、因果性を取り入れた要素は手術フェーズ認識のクロスドメイン一般化をどのように改善できるか。
- RQ2周波数領域の augmentation と因果抑制は、ベースラインの CLIP ベース手法と比べてドメインシフトに対する頑健性を共同で向上させるか。
- RQ3視覚言語アプローチは多模態の監視を活用して SurgVisDom の hard adaptation でドメン適応ベースラインを超えられるか。
- RQ4提案各要素の全体性能への寄与は是什么。
主な発見
- CauCLIP は SurgVisDom の hard adaptation ベンチマークにおいて、重み付き F1、非重み付き F1、グローバル F1、バランス精度の全指標で最先端の性能を達成した。
- Rand、SK、Parakeet、ResNet-50、ViT-B/16、SDA-CLIP など複数のベースラインをすべての報告指標で上回った。
- アブレーション研究では、因果性に基づく抑制(L_sup)と周波数領域拡張(L_aug)の双方が改善をもたらし、完全モデルが最良の性能を示した。
- 拡張と抑制の組み合わせは補完的な利点を生み出し、スタイル変動への頑健性を高め、因果的な手術意味論を強調する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。