[論文レビュー] SIP-SegNet: A Deep Convolutional Encoder-Decoder Network for Joint Semantic Segmentation and Extraction of Sclera, Iris and Pupil based on Periocular Region Suppression
SIP-SegNet は、適応的しきい値処理とファジィフィルタリングによる周図領域の抑制を通じて、制約のない眼画像において虹彩、網膜、瞳孔を共同でセグメンテーションする深層エンコーダデコーダネットワークである。CASIA データセットを用いた DnCNN ノイズ除去、CLAHE 増幅、および密接接続型完全畳み込みアーキテクチャを用いることで、それぞれ網膜、虹彩、瞳孔の平均 F1 スコアは 93.35、95.11、96.69 を達成した。
The current developments in the field of machine vision have opened new vistas towards deploying multimodal biometric recognition systems in various real-world applications. These systems have the ability to deal with the limitations of unimodal biometric systems which are vulnerable to spoofing, noise, non-universality and intra-class variations. In addition, the ocular traits among various biometric traits are preferably used in these recognition systems. Such systems possess high distinctiveness, permanence, and performance while, technologies based on other biometric traits (fingerprints, voice etc.) can be easily compromised. This work presents a novel deep learning framework called SIP-SegNet, which performs the joint semantic segmentation of ocular traits (sclera, iris and pupil) in unconstrained scenarios with greater accuracy. The acquired images under these scenarios exhibit purkinje reflexes, specular reflections, eye gaze, off-angle shots, low resolution, and various occlusions particularly by eyelids and eyelashes. To address these issues, SIP-SegNet begins with denoising the pristine image using denoising convolutional neural network (DnCNN), followed by reflection removal and image enhancement based on contrast limited adaptive histogram equalization (CLAHE). Our proposed framework then extracts the periocular information using adaptive thresholding and employs the fuzzy filtering technique to suppress this information. Finally, the semantic segmentation of sclera, iris and pupil is achieved using the densely connected fully convolutional encoder-decoder network. We used five CASIA datasets to evaluate the performance of SIP-SegNet based on various evaluation metrics. The simulation results validate the optimal segmentation of the proposed SIP-SegNet, with the mean f1 scores of 93.35, 95.11 and 96.69 for the sclera, iris and pupil classes respectively.
研究の動機と目的
- 反射、低解像度、まぶた・まつげの遮蔽、視線の変化といった、制約のない環境下における眼のセグメンテーションの課題に対処すること。
- 正確な眼の特徴のセグメンテーションを可能にすることで、マルチモーダルバイオメトリクスシステムの耐障害性と正確性を向上させること。
- 一般的なノイズ源であり誤分類の原因となる周図領域からの干渉を、適応的抑制技術によって低減すること。
- 網膜、虹彩、瞳孔の意味的セグメンテーションを高精度で同時に実行する統合的なディープラーニングフレームワークの開発すること。
提案手法
- 生眼画像のノイズ低減のため、ノイズ除去畳み込みニューラルネットワーク(DnCNN)を用いた前処理。
- 反射除去と画像強調のため、対照制限付き適応的ヒストограм等値化(CLAHE)の適用。
- 適応的しきい値処理を用いた周図領域の抽出により、注目領域を分離。
- ファジィフィルタリングによる周図情報の抑制により、セグメンテーション中の干渉を最小限に抑える。
- 密接接続型完全畳み込みエンコーダデコーダネットワーク(SIP-SegNet)を用いた網膜、虹彩、瞳孔の意味的セグメンテーション。
- 多様な撮影条件にわたる一般化を保証するため、5つの CASIA データセットを用いたエンドツーエンドのトレーニングと推論。
実験結果
リサーチクエスチョン
- RQ1制約のない眼画像において、人為的介入を最小限に抑え、深層学習フレームワークが網膜、虹彩、瞳孔を同時にかつ正確にセグメンテーションできるか。
- RQ2挑戦的な撮影条件下で、眼の構造物のセグメンテーション精度を向上させるために、周図領域の抑制はどの程度効果的か。
- RQ3DnCNN と CLAHE を用いた前処理は、眼のセグメンテーションネットワークの性能をどの程度向上させるか。
- RQ4ファジィフィルタリングによる周図領域抑制は、従来の前処理手法と比較して、セグメンテーションの F1 スコアにどのような差をもたらすか。
- RQ5SIP-SegNet は、低解像度、反射、遮蔽が生じる状況を含め、多様な実世界の眼画像データセットにおいて、どの程度の性能を示すか。
主な発見
- SIP-SegNet は、網膜セグメンテーションで平均 F1 スコア 93.35 を達成し、制約のない画像において高い正確性と再現率を示した。
- 虹彩セグメンテーションでは平均 F1 スコア 95.11 を達成し、反射や遮蔽に対して強い耐性を示した。
- 瞳孔セグメンテーションでは最高の平均 F1 スコア 96.69 を記録し、小規模でコントラストが低い領域において優れた性能を発揮した。
- DnCNN ノイズ除去と CLAHE 増幅の統合により、特にコントラストが低くノイズの多い画像において、セグメンテーション品質が顕著に向上した。
- 適応的しきい値処理とファジィフィルタリングによる周図領域の抑制により、誤検出が減少し、眼の構造物の境界の局所化が改善された。
- フレームワークは5つの CASIA データセットにおいて一貫した性能を示し、実世界のバイオメトリクス応用における一般化能力を裏付けた。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。