[논문 리뷰] Counterfactual Samples Synthesizing and Training for Robust Visual Question Answering
이 논문은 시각적 설명 가능성과 질문 민감성 능력을 향상시켜 시각질문응답(VQA) 모델의 성능을 향상시키기 위해 모델에 종속되지 않는 역설적 샘플 생성 및 훈련(CSST) 프레임워크를 제안한다. CSST는 중요한 시각적 객체나 질문 단어를 마스킹하여 역설적 샘플을 생성하고, 교차 엔트로피 손실과 대비 손실을 함께 사용하여 모델을 훈련함으로써 VQA-CP v2 및 GQA-OOD와 같은 분포 외 벤치마크에서 최고 성능을 달성한다.
Today's VQA models still tend to capture superficial linguistic correlations in the training set and fail to generalize to the test set with different QA distributions. To reduce these language biases, recent VQA works introduce an auxiliary question-only model to regularize the training of targeted VQA model, and achieve dominating performance on diagnostic benchmarks for out-of-distribution testing. However, due to complex model design, these ensemble-based methods are unable to equip themselves with two indispensable characteristics of an ideal VQA model: 1) Visual-explainable: The model should rely on the right visual regions when making decisions. 2) Question-sensitive: The model should be sensitive to the linguistic variations in questions. To this end, we propose a novel model-agnostic Counterfactual Samples Synthesizing and Training (CSST) strategy. After training with CSST, VQA models are forced to focus on all critical objects and words, which significantly improves both visual-explainable and question-sensitive abilities. Specifically, CSST is composed of two parts: Counterfactual Samples Synthesizing (CSS) and Counterfactual Samples Training (CST). CSS generates counterfactual samples by carefully masking critical objects in images or words in questions and assigning pseudo ground-truth answers. CST not only trains the VQA models with both complementary samples to predict respective ground-truth answers, but also urges the VQA models to further distinguish the original samples and superficially similar counterfactual ones. To facilitate the CST training, we propose two variants of supervised contrastive loss for VQA, and design an effective positive and negative sample selection mechanism based on CSS. Extensive experiments have shown the effectiveness of CSST. Particularly, by building on top of model LMH+SAR, we achieve record-breaking performance on all OOD benchmarks.
연구 동기 및 목표
- 훈련 데이터에서의 표면적인 언어적 상관관계에 의존하는 VQA 모델의 언어 편향 문제를 지속적으로 해결하기 위해.
- 시각적 설명 가능성과 질문 민감성 행동을 강제하여 예측이 올바른 시각적 및 언어적 단서에 기반하도록 모델의 강인성을 향상시키기 위해.
- 아키텍처 변경이나 앙상블 기반 설계가 필요 없는 모델에 종속되지 않는 훈련 전략을 개발하기 위해.
- 질문-답변 분포가 이동한 분포 외 테스트 세트에 대해 효과적으로 일반화할 수 있도록 VQA 모델을 가능하게 하기 위해.
- 원본과 역설적 샘플을 구분하도록 강제하는 데이터 증강 및 대비 학습 파이프라인을 설계하기 위해.
제안 방법
- 역설적 샘플 생성(CSS)은 중요한 시각적 객체나 질문의 핵심 단어를 마스킹하여 수정된 입력 기반의 가짜 정답을 할당함으로써 훈련 샘플을 생성한다.
- 역설적 샘플 훈련(CST)은 표준 교차 엔트로피 손실과 원본 및 역설적 샘플을 구분하는 대비 손실을 함께 사용하는 이중 목적의 손실을 사용한다.
- CSS 출력 기반으로 새로운 양성-음성 샘플 선택 메커니즘이 설계되어 대비 학습의 효과를 향상시킨다.
- VQA에 특화된 두 가지 변형된 감독형 대비 손실이 제안되어 원본 및 역설적 입력 간의 특징 분리 능력을 향상시킨다.
- 프레임워크는 아키텍처 변경 없이 기존 VQA 모델에 훈련 후 적용 가능하므로 모델에 종속되지 않으며 광범위하게 적용 가능하다.
- Grad-CAM과 POS 태깅을 사용하여 역설적 생성 중 마스킹 대상이 되는 중요한 시각적 객체와 단어를 식별한다.
실험 결과
연구 질문
- RQ1역설적 데이터 증강이 올바른 시각적 영역에 주의를 기울이게 하여 VQA 모델의 시각적 설명 가능성을 향상시킬 수 있는가?
- RQ2역설적 샘플로 훈련하면 언어적 변형 상황에서도 VQA 모델의 질문 민감성 행동이 향상되는가?
- RQ3모델에 종속되지 않는 훈련 전략이 분포 외 VQA 벤치마크에서 앙상블 기반 방법보다 성능이 뛰어나게 작용할 수 있는가?
- RQ4제안된 대비 손실은 원본과 유사하게 보이는 역설적 샘플을 효과적으로 구분하는 데 얼마나 효과적인가?
- RQ5CSST 프레임워크가 가장 잘 향상시키는 질문 유형과 가장 잘 향상시키지 못하는 질문 유형은 무엇이며, 그 이유는 무엇인가?
주요 결과
- CSST는 VQA-CP v2, VQA-CP v1, GQA-OOD를 포함한 모든 평가된 분포 외 벤치마크에서 최고 성능을 달성한다.
- VQA-CP v2에서 LMH-CSST 모델은 이전 최고 성능 모델들을 뛰어넘는 새로운 기록인 76.8%의 정확도를 기록한다.
- 신뢰도 향상(CI) 지표는 CSST 모델가 핵심 단어를 제거했을 때 신뢰도가 유의미하게 감소함을 보여주어 더 강한 질문 민감성을 나타낸다.
- 절단 실험 결과 CSS 및 CST 구성 요소가 모두 필수적임을 확인하였으며, 특히 대비 학습을 통한 CST 기여도가 성능 향상에 더 크다.
- 실패 분석 결과 이 방법은 '예/아니요' 및 '숫자' 질문에서 가장 어려움을 겪는 것으로 나타났으며, 이는 데이터셋 편향과 히وري스틱 예측의 용이성 때문일 것이다.
- 정성 분석 결과 CSST는 Grad-CAM 주의 맵을 중요한 시각적 객체와 단어에 맞추어 개선함으로써 시각적 설명 가능성의 타당성을 검증한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.