[論文レビュー] SAM-based instance segmentation models for the automation of structural damage detection
本論文は、自動モルタルクラック検出のための新しいSAMベースのインスタンスセグメンテーション手法を提案し、学習可能な自己生成プロンプターとプロンプトフリーのデコーダー統合を導入する。LoRAで微調整されたモデルは、全体で約3%、クラックでは約6%のSOTA性能を上回り、単眼カメラとホフ変換に基づく正射影変換を用いてレーザースキャニングと比較して10%以内の誤差で正確なクラック寸法推定を可能にする。
Automating visual inspection for capturing defects based on civil structures appearance is crucial due to its currently labour-intensive and time-consuming nature. An important aspect of automated inspection is image acquisition, which is rapid and cost-effective considering the pervasive developments in both software and hardware computing in recent years. Previous studies largely focused on concrete and asphalt, with less attention to masonry cracks. The latter also lacks publicly available datasets. In this paper, we first present a corresponding data set for instance segmentation with 1,300 annotated images (640 pixels x 640 pixels), named as MCrack1300, covering bricks, broken bricks, and cracks. We then test several leading algorithms for benchmarking, including the latest large-scale model, the prompt-based Segment Anything Model (SAM). We fine-tune the encoder using Low-Rank Adaptation (LoRA) and proposed two novel methods for automation of SAM execution. The first method involves abandoning the prompt encoder and connecting the SAM encoder to other decoders, while the second method introduces a learnable self-generating prompter. In order to ensure the seamless integration of the two proposed methods with SAM encoder section, we redesign the feature extractor. Both proposed methods exceed state-of-the-art performance, surpassing the best benchmark by approximately 3% for all classes and around 6% for cracks specifically. Based on successful detection, we propose a method based on a monocular camera and the Hough Line Transform to automatically transform images into orthographic projection maps. By incorporating known real sizes of brick units, we accurately estimate crack dimensions, with the results differing by less than 10% from those obtained by laser scanning. Overall, we address important research gaps in automated masonry crack detection and size estimation.
研究の動機と目的
- モルタルクラックインスタンスセグメンテーションのための公開データセットの不足に対処すること。
- 深層学習を用いてモルタル構造物の損傷を自動検出することに焦点を当て、クラックと破損ブリックを対象とする。
- 現実世界の構造物点検において、Segment Anything Model (SAM) のプロンプト依存性の制限を克服すること。
- 幾何的変換を用いて単眼画像から正確なクラック寸法推定を可能とすること。
- 都市インfra構造物の自動視覚点検のためのスケーラブルで低コストなソリューションを開発すること。
提案手法
- ブリック、破損ブリック、クラックの画像1,300枚(640×640 px)をアノテートした新しいデータセット、MCrack1300を提案。
- 効率的なパrameter適応のため、LoRAを用いてSAMエンコーダーを微調整。
- プロンプトエンコーダーを代替デコーダーに置き換えることで、プロンプトフリーな手法を導入。
- 自律的かつ効果的なプロンプトを生成できる学習可能な自己生成プロンプターを開発。
- 両手法とSAMエンコーダーのシームレスな統合を可能とするために、特徴抽出器を再設計。
- 実世界の画像から正射影マップを生成するために、単眼カメラ画像とホフ変換を組み合わせた手法を適用。
実験結果
リサーチクエスチョン
- RQ1標準的なプロンプトベース推論と比較して、微調整済みのプロンプトフリーなSAMの適応は、モルタルクラックのインスタンスセグメンテーション性能を向上させることができるか?
- RQ2学習可能な自己生成プロンプターは、自動構造物損傷検出において人為的プロンプトの代わりに効果的に機能するか?
- RQ3既知のブリック寸法と組み合わせた単眼カメラベースの正射影変換によって、どの程度正確なクラック長さ推定が可能か?
- RQ4すべてのクラックおよびモルタル欠損欠陊クラスにおいて、提案手法はSOTAモデルと比較してmAPでどの程度優れているか?
- RQ5レーザースキャニングの基準と比較して、クラック寸法推定誤差が10%未満に抑えられるか?
主な発見
- 提案されたプロンプトフリー手法は、すべての欠陊クラスにおいて、最良のベンチマークモデルと比較して平均平均精度(mAP)で3%の向上を達成した。
- クラック固有の検出において、mAPでSOTAを約6%上回った。
- 学習可能な自己生成プロンプターは、手動によるプロンプト設計を必要とせず、安定した性能を示した。
- 単眼カメラとホフ変換を用いたクラック長さ推定手法は、レーザースキャニング測定値と比較して10%未満の偏差を示した。
- MCrack1300データセットは、1,300枚の高解像度で完全にアノテートされた画像を備えており、モルタルクラックインスタンスセグメンテーションの新しいベンチマークを提供する。
- 再設計された特徴抽出器の統合により、両手法とも安定的かつ高精度な推論が可能になった。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。