[논문 리뷰] CauCLIP: Bridging the Sim-to-Real Gap in Surgical Video Understanding via Causality-Inspired Vision-Language Modeling
CauCLIP은 목표 도메인 데이터 없이 시뮬레이터-실제(domain) 간의 이동에 일반화되는 강건한 수술 단계 인식을 위한 인과관계 가이드 기반 CLIP 프레임워크를 제시하며, 주파수 기반 증강과 인과 차단 손실을 활용합니다.
Surgical phase recognition is a critical component for context-aware decision support in intelligent operating rooms, yet training robust models is hindered by limited annotated clinical videos and large domain gaps between synthetic and real surgical data. To address this, we propose CauCLIP, a causality-inspired vision-language framework that leverages CLIP to learn domain-invariant representations for surgical phase recognition without access to target domain data. Our approach integrates a frequency-based augmentation strategy to perturb domain-specific attributes while preserving semantic structures, and a causal suppression loss that mitigates non-causal biases and reinforces causal surgical features. These components are combined in a unified training framework that enables the model to focus on stable causal factors underlying surgical workflows. Experiments on the SurgVisDom hard adaptation benchmark demonstrate that our method substantially outperforms all competing approaches, highlighting the effectiveness of causality-guided vision-language models for domain-generalizable surgical video understanding.
연구 동기 및 목표
- 실제 비디오에 대한 제한된 주석으로 인한 시뮬레이터-실제 도메인 격차를 해결한다.
- CLIP를 기반으로 한 도메인 불변 표현을 학습하기 위한 인과관계에서 영감을 받은 컴퓨터 비전-언어 프레임워크를 개발한다.
- 주파수 기반 증강과 인과 차단 손실을 도입하여 비인과적 편향을 줄인다.
제안 방법
- 수술 단계 인식을 위한 영상-텍스트 정렬을 CLIP 기반의 영상-텍스트 매칭 작업으로 구축한다.
- 고주파수의 비인과적 단서를 교란시키되 의미를 보존하는 주파수 도메인 증강을 도입한다.
- 원래 표현과 주파수 증강 표현 간의 유사성을 강제하고 비인과적 특성과의 상관을 제거하는 인과 차단 모듈을 구현한다.
- augmented 뷰 간 의미 일관성을 보존하기 위한 증강 정렬 손실을 추가한다.
- 원래의 CLIP 손실, 증강 정렬 손실, 차단 손실을 합친 총 손실로 학습한다.
실험 결과
연구 질문
- RQ1대상 도메인 접근 없이 인과관계에서 영감을 받은 구성 요소가 수술 단계 인식의 교차 도메인 일반화를 어떻게 개선할 수 있는가?
- RQ2주파수 도메인 증강과 인과 차단이 기저 CLIP 기반 방법 대비 도메인 이동성에 대한 강건성을 공동으로 향상시키는가?
- RQ3비전-언어 접근법이 SurgVisDom의 어려운 적응에서 도메인 적응 기반 기준을 넘어설 수 있는가?
- RQ4제안된 각 구성 요소가 전체 성능에 기여하는 바는 무엇인가?
주요 결과
- CauCLIP은 가중치 F1, 비가중치 F1, 글로벌 F1, 균형 정확도에서 SurgVisDom 어려운 적응 벤치마크에서 최첨단 성능을 달성한다.
- Rand, SK, Parakeet, ResNet-50, ViT-B/16, SDA-CLIP를 포함한 다수의 기준선보다 모든 보고 지표에서 우위이다.
- 인과관계에서 영감을 받은 차단(L_sup)과 주파수 도메인 증강(L_aug) 모두 이점을 제공하며, 전체 모델이 최적의 성능을 보인다.
- 증강과 차단의 결합은 보강 효과를 보이며 스타일 변 Variation에 대한 강건성을 높이고 인과적 수술 시맨틱스를 강조한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.