Skip to main content
QUICK REVIEW

[論文レビュー] High-Resolution Optical Flow from 1D Attention and Correlation

Haofei Xu, Jiaolong Yang|arXiv (Cornell University)|Apr 28, 2021
Advanced Vision and Imaging参考文献 44被引用数 7
ひとこと要約

本稿では、高解像度オプティカルフロー推定のための、1次元アテンションと相関に基づくコストボリューム構築手法Flow1Dを提案する。2次元の対応マッチングを直交する1次元アテンションと相関演算に分解することで、計算複雑度をO(H²W²)からO(HW(H+W))に低減し、RAFTに比べて6倍のメモリ削減を実現しながら、8K解像度の画像処理を効率的に行える。SintelおよびKITTIベンチマークでも競争力のある精度を維持している。

ABSTRACT

Optical flow is inherently a 2D search problem, and thus the computational complexity grows quadratically with respect to the search window, making large displacements matching infeasible for high-resolution images. In this paper, we take inspiration from Transformers and propose a new method for high-resolution optical flow estimation with significantly less computation. Specifically, a 1D attention operation is first applied in the vertical direction of the target image, and then a simple 1D correlation in the horizontal direction of the attended image is able to achieve 2D correspondence modeling effect. The directions of attention and correlation can also be exchanged, resulting in two 3D cost volumes that are concatenated for optical flow estimation. The novel 1D formulation empowers our method to scale to very high-resolution input images while maintaining competitive performance. Extensive experiments on Sintel, KITTI and real-world 4K ($2160 imes 3840$) resolution images demonstrated the effectiveness and superiority of our proposed method. Code and models are available at \url{https://github.com/haofeixu/flow1d}.

研究の動機と目的

  • RAFTのような最先端のオプティカルフロー手法における4次元コストボリュームの2次関数的計算複雑度を解消し、高解像度入力へのスケーラビリティを向上させること。
  • 4Kや8Kなどの高解像度画像を処理する際の、従来の4次元コストボリューム構築の非効率性とメモリ制限を克服すること。
  • 競争力のある精度を維持しながら、超高解像度動画におけるリアルタイム推論を可能にする、コンactなコストボリューム定式化を開発すること。
  • 消費デバイス向けの高解像度動画における実用的なオプティカルフロー推定を実現し、キーマッチングモデリングを損なわず、メモリと計算量を削減すること。

提案手法

  • ターゲット画像特徴マップに縦方向に1次元アテンションを適用し、行間をまたがる長距離依存性を伝搬する。
  • アテンション処理済み特徴に対して横方向に1次元相関を適用し、水平方向の対応をモデル化する3次元コストボリュームを構築する。これにより2次元マッチングを効果的にシミュレートする。
  • アテンションと相関の方向を逆転させることで、2番目の3次元コストボリュームを生成し、最初のものと連結することで完全な2次元対応マッピングを実現する。
  • サイズH×W×WおよびH×W×Hの2つの3次元コストボリュームを構築することで、従来の4次元相関と比較して、全体の複雑度をO(H²W²)からO(HW(H+W))に低減する。
  • 得られたコンactなコストボリュームを、エンドツーエンドのオプティカルフロー回帰用のフローヘッドネットワークの入力として使用する。
  • Transformerにインspiredされた自己アテンションおよびクロスアテンション機構を活用し、1次元定式化における特徴伝搬とマッチング精度を向上させる。
Figure 1 : Optical flow factorization. We factorize the 2D optical flow with 1D attention and correlation in orthogonal directions to achieve large displacements search on high-resolution images. Specifically, for the correspondence ( red point) of the blue point, we first perform a 1D vertical atte
Figure 1 : Optical flow factorization. We factorize the 2D optical flow with 1D attention and correlation in orthogonal directions to achieve large displacements search on high-resolution images. Specifically, for the correspondence ( red point) of the blue point, we first perform a 1D vertical atte

実験結果

リサーチクエスチョン

  • RQ12次元オプティカルフロー対応マッピングを、顕著な精度損失なしに直交する1次元演算に分解できるか?
  • RQ21次元アテンションと相関は、計算コストを著しく低減しながらも、標準ベンチマークで競争力のある性能を達成できるか?
  • RQ3既存の4次元コストボリューム手法と比較して、この手法はどれほど超高解像度画像(例:4Kおよび8K)にスケーリングできるか?
  • RQ4高解像度入力において、RAFTのような最先端モデルと比較して、本手法のメモリ効率および推論速度はどの程度優れているか?

主な発見

  • Sintelテストスプリットでは2番目に高い性能を達成し、最終的なエンドポイント誤差は1.25であり、RAFT(1.27)に次いで、PWC-Net+(2.34)を上回った。
  • KITTI 2015ベンチマークではF1-allスコア6.27を達成し、PWC-Net+(7.72)を上回り、MaskFlowNet(6.10)に近いが、依然としてRAFT(5.10)にわずかに劣った。
  • 1080p解像度では、RAFTに比べて6倍のメモリ消費量を抑え、コストボリューム計算の削減により著しく高速な推論を実現した。
  • Flow1Dは、4K(2160×3840)解像度の画像を5.8GBのメモリで処理でき、RAFTはメモリ不足により失敗した。
  • 8K解像度(4320×7680)にも21.81GBのメモリ消費量でスケーリング可能であり、現在の最先端手法をはるかに超えるスケーラビリティを示した。
  • DAVIS 1080pおよび4Kデータセットにおける可視化比較では、Flow1DはRAFに匹敵する結果を生成し、目立つアーティファクトが少なく、優れた構造的一致性を示した。
Figure 2 : Overview of our framework. Given a pair of source and target images, we first extract $8\times$ downsampled features with a shared backbone network. The features are then used to construct two 3D cost volumes with vertical attention, horizontal correlation and horizontal attention, vertic
Figure 2 : Overview of our framework. Given a pair of source and target images, we first extract $8\times$ downsampled features with a shared backbone network. The features are then used to construct two 3D cost volumes with vertical attention, horizontal correlation and horizontal attention, vertic

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。