Skip to main content
QUICK REVIEW

[論文レビュー] Some open questions on morphological operators and representations in the deep learning era

Jesús Angulo|arXiv (Cornell University)|May 4, 2021
Topological and Geometric Data Analysis参考文献 95被引用数 8
ひとこと要約

本稿は、現代のAIパラダイムにおけるマトロジカルな演算子と表現の再考を通じて、数学的モルフォロジーとディープラーニングを統合する研究アジェンダを提唱する。ニューラルネットワーク、NLP、プログラム合成などのディープラーニング技術を活用し、自動的にモルフォロジカル演算子、構造要素、モルフォロジカルプログラムを学習することを目指しており、理論的モルフォロジーとエンドツーエンド微分可能なアーキテクチャを統合することで、AIシステムの解釈可能性と設計を向上させることを目的としている。

ABSTRACT

During recent years, the renaissance of neural networks as the major machine learning paradigm and more specifically, the confirmation that deep learning techniques provide state-of-the-art results for most of computer vision tasks has been shaking up traditional research in image processing. The same can be said for research in communities working on applied harmonic analysis, information geometry, variational methods, etc. For many researchers, this is viewed as an existential threat. On the one hand, research funding agencies privilege mainstream approaches especially when these are unquestionably suitable for solving real problems and for making progress on artificial intelligence. On the other hand, successful publishing of research in our communities is becoming almost exclusively based on a quantitative improvement of the accuracy of any benchmark task. As most of my colleagues sharing this research field, I am confronted with the dilemma of continuing to invest my time and intellectual effort on mathematical morphology as my driving force for research, or simply focussing on how to use deep learning and contributing to it. The solution is not obvious to any of us since our research is not fundamental, it is just oriented to solve challenging problems, which can be more or less theoretical. Certainly, it would be foolish for anyone to claim that deep learning is insignificant or to think that one's favourite image processing domain is productive enough to ignore the state-of-the-art. I fully understand that the labs and leading people in image processing communities have been shifting their research to almost exclusively focus on deep learning techniques. My own position is different: I do think there is room for progress on mathematically grounded image processing branches, under the condition that these are rethought in a broader sense from the deep learning paradigm. Indeed, I firmly believe that the convergence between mathematical morphology and the computation methods which gravitate around deep learning (fully connected networks, convolutional neural networks, residual neural networks, recurrent neural networks, etc.) is worthwhile. The goal of this talk is to discuss my personal vision regarding these potential interactions. Without any pretension of being exhaustive, I want to address it with a series of open questions, covering a wide range of specificities of morphological operators and representations, which could be tackled and revisited under the paradigm of deep learning. An expected benefit of such convergence between morphology and deep learning is a cross-fertilization of concepts and techniques between both fields. In addition, I think the future answer to some of these questions can provide some insight on understanding, interpreting and simplifying deep learning networks.

研究の動機と目的

  • ニューラルネットワークの時代において、モルフォロジーが陳腐化していると見なされるのを防ぐために、ディープラーニングフレームワークに数学的モルフォロジーを埋め込むことによる、その活性化を図ること。
  • 数学的に根拠のある一方で、現代のディープラーニングパイプラインと互換性を持つモルフォロジカル演算子および表現の設計という課題に取り組むこと。
  • 教師あり学習、遺伝的アルゴリズム、PAC学習を用いて、入出力画像ペアからデータ駆動でモルフォロジカル演算子および合成を自動的に発見すること。
  • 自然言語処理技術を活用し、モルフォロジカルプログラムを「テキスト」としてモデル化することで、単語埋め込みおよびプログラム埋め込みの学習を可能とすること。
  • 解釈可能性、一般化、アーキテクチャの革新を高めるために、体系的かつ理論的根拠に基づいたモルフォロジカルAIのアプローチを構築すること。

提案手法

  • 画像ペアを用いた確率的勾配降下法による畳み込みニューラルネットワークの訓練を通じて、モルフォロジカル演算子(例:エロージョン、ディレイション、オープニング、クロージング)をディープラーニングで学習する。
  • モルフォロジカルプログラムを「ディレイションとディスク」や「ラインによるオープニング」などの操作の系列としてモデル化し、word2vec や context2vec などのNLP技術を用いて、演算子および構造要素の埋め込みを学習する。
  • 演算子と構造要素のための構造的トークンを有するモルフォロジカル言語を定義し、意味的表現力と学習可能性のバランスを取るために粒度を最適化する。
  • プログラム合成と組み合わせ最適化を用いてモルフォロジカルプログラム空間を探索し、ディープラーニングを用いて候補プログラムの順序付けと単純化を行う。
  • モルフォロジカルパーセプトロンにインspiredされた、微分可能なモジュールとしてのモルフォロジカルニューラルネットワークとアソシエイティブメモリを、ディープアーキテクチャに統合する。
  • ラティス理論、チョケット容量、ハミルトニアン・ヤコビPDEなど理論的基盤を活用し、微分可能なモルフォロジカル層の設計をガイドし、数学的整合性を保証する。

実験結果

リサーチクエスチョン

  • RQ1どのようにして、勾配降下法によるエンドツーエンド学習を通じて、モルフォロジカル演算子を効果的にパラメータ化し、学習することができるか?
  • RQ2NLPベースの学習に適した形式的言語として、モルフォロジカルプログラムをどのように表現すべきか。また、演算子と構造要素のコンponentsをどのようにトークン化すべきか?
  • RQ3大規模な熟練者によるモルフォロジカルプログラムのコーパスを抽出し、下流のプログラム生成タスクに活用できる意味のある埋め込み(例:operator2vec)を事前学習することができるか?
  • RQ4ディープラーニングと組み合わせ最適化をどのように統合することで、大規模なモルフォロジカルプログラム空間を探索し、効率的で解釈可能な合成を発見できるか?
  • RQ5ラティス理論、イドムポテンス代数、多スケール半群などの数学的モルフォロジーの理論的仕組みを、微分可能なディープラーニングアーキテクチャにどのように埋め込むことができるか。その結果、解釈可能性と一般化がどのように向上するか?

主な発見

  • ディープラーニングを用いることで、反調和平均を漸近的近似として用いることで、エロージョンやディレイションなどのモルフォロジカル演算子を学習可能となり、構造要素および合成のエンドツーエンド学習が可能になる。
  • 既存の手法(例:[47])は、CNNがオープニングとクロージングの合成を通じてTV正則化を近似できることを示しており、モルフォロジカルパイプラインの学習の可能性を裏付けている。
  • モルフォロジカルプログラムにNLP技術を適用することは有望であるが、有効な事前学習に使用可能な十分に大きなラベル付きコーパスが現在不足しているため、制限がある。
  • ディープラーニングを用いたプログラム合成は、短いドメイン特化のモルフォロジカルプログラムにしか対応できず、スケーラブルな探索およびモジュラー分解戦略の導入が求められる。
  • モルフォロジーの理論的基盤(例:ラティスに基づく定式化、チョケット容量、トロピカル幾何)は、学習されたモルフォロジカルモデルのロバスト性と解釈可能性を高める強固な数学的基盤を提供する。
  • 数学的モルフォロジーとディープラーニングの間には、一致するだけでなく、相互に補い合う関係が成立し、形状に基づく解釈可能な演算子を通じてニューラルネットワークの挙動のより深い理解が可能になる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。