[論文レビュー] Both Style and Fog Matter: Cumulative Domain Adaptation for Semantic Foggy Scene Understanding
本稿では、セマンティックな曇りガラスシーン理解(SFSU)におけるスタイル要因と曇り要因を分離する累積的ドメイン適応フレームワーク、CuDA-Netを提案する。この手法は、中間ドメインを介して、クリアから曇りドメインへの段階的適応を可能にし、3つのベンチマークで最先端の性能を達成する。また、新規の累積損失関数を用いてスタイル、曇り、および二重要因を同時に分離することで、雨や雪のシーンへの一般化も可能となる。
Although considerable progress has been made in semantic scene understanding under clear weather, it is still a tough problem under adverse weather conditions, such as dense fog, due to the uncertainty caused by imperfect observations. Besides, difficulties in collecting and labeling foggy images hinder the progress of this field. Considering the success in semantic scene understanding under clear weather, we think it is reasonable to transfer knowledge learned from clear images to the foggy domain. As such, the problem becomes to bridge the domain gap between clear images and foggy images. Unlike previous methods that mainly focus on closing the domain gap caused by fog -- defogging the foggy images or fogging the clear images, we propose to alleviate the domain gap by considering fog influence and style variation simultaneously. The motivation is based on our finding that the style-related gap and the fog-related gap can be divided and closed respectively, by adding an intermediate domain. Thus, we propose a new pipeline to cumulatively adapt style, fog and the dual-factor (style and fog). Specifically, we devise a unified framework to disentangle the style factor and the fog factor separately, and then the dual-factor from images in different domains. Furthermore, we collaborate the disentanglement of three factors with a novel cumulative loss to thoroughly disentangle these three factors. Our method achieves the state-of-the-art performance on three benchmarks and shows generalization ability in rainy and snowy scenes.
研究の動機と目的
- 悪天候下における視界の低下とラベル付き曇りガラスデータの不足という課題に、セマンティックな曇りガラスシーン理解(SFSU)を対象とする。
- 従来のドメイン適応手法がドメインギャップを単一の問題として扱い、スタイルの変動と曇りによる劣化の違いを無視するという限界を克服する。
- スタイルと曇り要因を明示的に分離し、別々に適応する手法を提案する。中間ドメインを用いて二重要因ドメインギャップを分解する。
- スタイル、曇り、および二重要因適応の累積的関係をモデル化することで、SFSUにおける性能と一般化能力を向上させる。
提案手法
- ソース(クリア)、中間(クリア)、ターゲット(曇り)ドメインからの画像からスタイル要因と曇り要因を分離する統一フレームワークを導入する。
- 段階的適応を強制するための新規の累積損失関数を採用:まずソースから中間ドメインへのスタイル適応、次に中間からターゲットへの曇り適応、最後にすべてのドメイン間での二重要因適応。
- 中間ドメイン(例:Zurichのクリア画像)を用いて、混合ドメインギャップをスタイルギャップ(ソース–中間)と曇りギャップ(中間–ターゲット)に分離する。
- 分離ネットワークを用いてコンテンツ、スタイル、曇り要因を分離し、各要因を独立して適応可能にする。
- 最適化の安定化と分離品質の向上を目的に、ハイパーパrameter Tを用いたサイクルトレーニングを適用する。
- 累積損失におけるドメイン差違の測定にL2距離を最適な指標として使用し、アブレーションスタディで妥当性を検証した。
実験結果
リサーチクエスチョン
- RQ1セマンティックな曇りガラスシーン理解におけるドメインギャップは、スタイル要因と曇り関連要因に効果的に分解可能か?
- RQ2まずスタイルに適応し、次に曇りに適応し、最後に両者を同時に適応する累積的適応は、直接的ドメイン適応よりも性能を向上させるか?
- RQ3手動で選択された中間ドメイン(クリア画像)は、CNNベースの選択手法に比べて優れた性能を発揮するか?
- RQ4提案手法は、雨や雪といった他の悪天候条件に対しても一般化可能か?
- RQ5累積損失はスタイル、曇り、および二重要因の分離に有効であるか?また、ハイパーパrameterにどれほど感度を示すか?
主な発見
- CuDA-NetはFoggy Zurichテストセットで43.06 mIoUを達成し、先行研究の最先端手法を大きく上回る性能を示した。
- 完全なパイプライン(F_s→m + F_m→t + F_s→t)は、直接適応(F_s→t)に比べて2.28 mIoUの向上を示し、二重要因適応の必要性を裏付けた。
- T=2のサイクルトレーニングは2.72 mIoUの性能向上をもたらし、最適化の安定化に有効であることを示した。
- L2距離指標は累積損失において最良の性能を発揮し、他の距離指標よりもアブレーションスタディで優れた結果を示した。
- 手動による中間ドメイン選択(Clear Zurich)は、CNNベースの選択を上回り、人間が検証したクリア画像の価値を確認した。
- 本手法は雨や雪のシーンに対しても良好な一般化性能を示し、ACDC(雨)では48.5 mIoU、ACDC(雪)では47.2 mIoUを達成し、直接適応を上回った。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。