Skip to main content
QUICK REVIEW

[논문 리뷰] F3Net: Fusion, Feedback and Focus for Salient Object Detection

Jun Wei, Shuhui Wang|arXiv (Cornell University)|2019. 11. 26.
Visual Attention and Saliency Detection참고 문헌 35인용 수 113
한 줄 요약

F3 Net은 교차 특징 모듈을 도입하여 선택적 다층 특징 융합, 다중 피드백 디코더를 통한 반복 정제, 하드 픽셀을 강조하기 위한 픽셀 위치 인식 손실을 도입해 다섯 개 데이터셋에서 최첨단 Salient Object Detection 성능을 달성한다.

ABSTRACT

Most of existing salient object detection models have achieved great progress by aggregating multi-level features extracted from convolutional neural networks. However, because of the different receptive fields of different convolutional layers, there exists big differences between features generated by these layers. Common feature fusion strategies (addition or concatenation) ignore these differences and may cause suboptimal solutions. In this paper, we propose the F3Net to solve above problem, which mainly consists of cross feature module (CFM) and cascaded feedback decoder (CFD) trained by minimizing a new pixel position aware loss (PPA). Specifically, CFM aims to selectively aggregate multi-level features. Different from addition and concatenation, CFM adaptively selects complementary components from input features before fusion, which can effectively avoid introducing too much redundant information that may destroy the original features. Besides, CFD adopts a multi-stage feedback mechanism, where features closed to supervision will be introduced to the output of previous layers to supplement them and eliminate the differences between features. These refined features will go through multiple similar iterations before generating the final saliency maps. Furthermore, different from binary cross entropy, the proposed PPA loss doesn't treat pixels equally, which can synthesize the local structure information of a pixel to guide the network to focus more on local details. Hard pixels from boundaries or error-prone parts will be given more attention to emphasize their importance. F3Net is able to segment salient object regions accurately and provide clear local details. Comprehensive experiments on five benchmark datasets demonstrate that F3Net outperforms state-of-the-art approaches on six evaluation metrics.

연구 동기 및 목표

  • 다중 수준 CNN 특징 간의 차이로 인한 불일치를 완화하여 융합 품질을 향상시키려 한다(다른 수용 영역).
  • 중복 정보를 억제하면서 보완적 세부 정보를 보존하는 선택적 융합 메커니즘을 도입한다.
  • 경계선과 국부적 디테일을 선명하게 하기 위해 반복적, 계단식 피드백으로 다층 특징을 정제한다.
  • 하드 픽셀과 경계 및 세부가 풍부한 영역을 강조하기 위해 로컬 구조 맥락에 가중치를 두는 손실을 도입한다.
  • MLS를 포함한 다중 벤치마크 데이터셋에서 포괄적 어브레이션을 통해 최첨단 성능을 입증한다.

제안 방법

  • 저수준 및 고수준 특징을 요소 곱으로 융합하고 원래 특징을 다시 더해 표현을 정제하는 Cross Feature Module (CFM)을 제안한다.
  • CFM을 통해 하향식 피드백과 함께 다중 서브-디코더가 바닥에서 위쪽으로 합치고 이전 계층으로 피드백을 받아 그리드 형태의 정제 네트워크를 형성하는 Cascaded Feedback Decoder (CFD)를 개발한다.
  • Hard 픽셀을 강조하고 로컬 구조를 보존하기 위해 위치 의존 가중 BCE와 가중 IoU 손실을 결합한 Pixel Position Aware (PPA) 손실을 도입한다.
  • 최적화를 안정화하고 학습을 향상시키기 위해 CFD 단계 전반에 걸친 MLS(다중 수준 감독)로 학습한다.
  • 백본으로 ResNet-50를 사용하고 점진적으로 정제된 디코딩 경로를 통해 포스트 프로세싱 없이 최종 saliency 맵을 생성한다.
  • 다섯 개 데이터셋에서 평가하고 CFM, CFD, PPA의 기여를 보여주는 어브레이션을 수행한다.

실험 결과

연구 질문

  • RQ1샤펜시 맵에서 redundancy를 줄이고 의미적 및 공간적 세부 정보를 보존하기 위해 다층 CNN 특징을 더 효과적으로 융합할 수 있는 방법은 무엇인가?
  • RQ2피드백 기반 디코딩 방식이 경계 정밀도와 로컬 구조를 회복하기 위해 특징을 반복적으로 정제할 수 있는가?
  • RQ3픽셀 위치를 고려한 손실이 표준 BCE나 IoU 손실에 비해 하드 픽셀과 경계 영역의 탐지 개선에 기여하는가?
  • RQ4MLS, CFM, CFD를 통합하는 실험이 벤치마크에서 표준 saliency 지표에 미치는 경험적 영향은 무엇인가?

주요 결과

  • F3 Net은 다섯 개 데이터셋에서 여섯 가지 지표에서 최신 방법보다 우수한 성능을 달성한다.
  • CFM은 배경 잡음을 효과적으로 억제하고 선택적 융합을 통해 고수준 경계를 선명하게 한다.
  • CFD는 정제된 특징을 이전 계층으로 피드백하여 반복적인 정제를 가능하게 하여 주목도 맵의 품질을 향상시킨다.
  • PPA 손실은 하드 픽셀을 강조하고 로컬 구조를 통합하여 경계 및 디테일 표현을 개선한다.
  • 어브레이션 연구에서 CFM, CFD, MLS, PPA가 함께 최상의 결과에 기여함을 보여준다.
  • 모델은 포스트 프로세싱 없이도 더 뚜렷한 주목도 맵과 배경 잡음 감소를 달성하며 견고한 성능을 제공한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.