Skip to main content
QUICK REVIEW

[論文レビュー] Deep Convolutional Neural Networks in the Face of Caricature: Identity and Image Revealed

Matthew Q. Hill, Connor J. Parde|arXiv (Cornell University)|Dec 28, 2018
Face recognition and analysis参考文献 57被引用数 13
ひとこと要約

この論文は、深層畳み込みニューラルネットワーク(DCNNs)が、高次元の顔空間内において顔の同一性と撮影要因(視点、照明、性別など)をどのように組織化しているかを調査する。t-SNE可視化と線形分類を用いて、同一性は性別のもとで階層的にネストされ、照明は同一性のもとで、視点は照明のもとでネストされていることが示され、コマーシャル(風刺的顔)は特徴を強調することで同一性の識別を向上させ、外見の変動に対する感受性を低下させる。

ABSTRACT

Real-world face recognition requires an ability to perceive the unique features of an individual face across multiple, variable images. The primate visual system solves the problem of image invariance using cascades of neurons that convert images of faces into categorical representations of facial identity. Deep convolutional neural networks (DCNNs) also create generalizable face representations, but with cascades of simulated neurons. DCNN representations can be examined in a multidimensional "face space", with identities and image parameters quantified via their projections onto the axes that define the space. We examined the organization of viewpoint, illumination, gender, and identity in this space. We show that the network creates a highly organized, hierarchically nested, face similarity structure in which information about face identity and imaging characteristics coexist. Natural image variation is accommodated in this hierarchy, with face identity nested under gender, illumination nested under identity, and viewpoint nested under illumination. To examine identity, we caricatured faces and found that network identification accuracy increased with caricature level, and--mimicking human perception--a caricatured distortion of a face "resembled" its veridical counterpart. Caricatures improved performance by moving the identity away from other identities in the face space and minimizing the effects of illumination and viewpoint. Deep networks produce face representations that solve long-standing computational problems in generalized face recognition. They also provide a unitary theoretical framework for reconciling decades of behavioral and neural results that emphasized either the image or the object/face in representations, without understanding how a neural code could seamlessly accommodate both.

研究の動機と目的

  • 深層畳み込みニューラルネットワーク(DCNNs)が顔の同一性および視点、照明、性別などの撮影変動をどのように表現しているかを理解すること。
  • DCNNsが顔の同一性の不変性と画像レベルの変動を同時に維持できるかどうかを検証し、顔認識分野における長年の論争を解消すること。
  • コマーシャルのDCNNにおける同一性表現に与える影響を調査し、人間の顔の類似度認識に類似したモデルを構築すること。
  • DCNNにおける顔空間の階層的組織構造が、霊長類視覚系や行動データで観察された原則を反映しているかどうかを検証すること。
  • 深層学習表現を用いて、オブジェクト中心モデルと画像ベースモデルを統合する理論的枠組みを提供すること。

提案手法

  • 顔識別を目的とした2つのDCNNを訓練した:ネットワークA(UniverseデータセットにCrystal Lossを適用したResNet-101)とネットワークB(CASIA-WebFaceデータセットに15層CNNを適用)。
  • すべての刺激(顔のモーフィング顔を含む)に対して、512次元の直前層特徴を抽出し、同一性記述子として使用した。
  • t-SNE(Barnes-Hut近似、θ=0.5、perplexity=30および100)を用いて、高次元顔空間を2次元に可視化し、角距離を保持した。
  • 線形判別分析(LDA)を用いて性別および照明を分類し、Moore-Penrose一般化逆行列を用いた線形回帰で視点を予測した。
  • 分類結果の統計的有意性を評価するため、パーミュテーション検定(n=1000)を実施し、すべての変数でp<.001とした。
  • 3Dレーザースキャンを用いてモーフィング刺激を生成し、顔の同一性強度(s)を操作した。s>1でコマーシャル、0<s<1でアンチ・コマーシャルが得られた。

実験結果

リサーチクエスチョン

  • RQ1DCNNの深層顔認識空間内において、同一性、性別、照明、視点はどのように組織化されているか?
  • RQ2コマーシャルの強度を高めることでDCNNにおける同一性認識精度が向上するか?また、これは人間の顔の類似度認識と一致するか?
  • RQ3DCNNは視点や照明の変動に対しても、どの程度同一性の不変性を維持できるか?
  • RQ4撮影要因(例:性別のもとでの同一性、同一性のもとでの照明)の階層的ネスト構造を顔空間で定量化および可視化できるか?
  • RQ5コマーシャルは顔空間内での同一性表現の分離を高め、照明や視点による干渉を軽減するか?

主な発見

  • DCNNの顔空間は階層的かつネストされた構造を示しており、同一性は性別のもとで、照明は同一性のもとで、視点は照明のもとでネストされている。
  • コマーシャル顔はDCNNの同一性認識精度を向上させ、コマーシャルの強度が高くなるほど性能が向上した。
  • コマーシャルは同一性の特徴を強調することで、顔空間内での他の同一性からの距離を拡大し、照明や視点による干渉を低減した。
  • 512次元特徴に対するLDAを用いた性別および照明の線形分類は統計的に有意(p<.001)であり、真の値とパーミュテーション検定からのノイズ分布との重複はなかった。
  • Moore-Penrose一般化逆行列を用いた線形回帰による視点予測も有意(p<.001)であり、視点情報が表現に符号化されていることが確認された。
  • 結果は2つの異なるDCNNアーキテクチャ(ネットワークAおよびネットワークB)で一貫しており、顔空間組織の堅牢性が裏付けられた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。