[論文レビュー] Investigation into respiratory sound classification for an imbalanced data set using hybrid LSTM-KAN architectures
論文は、偏りのある呼吸音を分類するために焦点損失、データ拡張、SMOTEを用いたハイブリッド LSTM-KAN モデルを提案し、COPD dominant な6クラスデータセットで 94.6% の精度と macro F1 0.703 を達成します。
Respiratory sounds captured via auscultation contain critical clues for diagnosing pulmonary conditions. Automated classification of these sounds faces challenges due to subtle acoustic differences and severe class imbalance in clinical datasets. This study investigates respiratory sound classification with a focus on mitigating pronounced class imbalance. We propose a hybrid deep learning model that combines a Long Short-Term Memory (LSTM) network for sequential feature encoding with a Kolmogorov-Arnold Network (KAN) for classification. The model is integrated with a comprehensive feature extraction pipeline and targeted imbalance mitigation strategies. Experiments were conducted on a public respiratory sound database comprising six classes with a highly skewed distribution. Techniques such as focal loss, class-specific data augmentation, and Synthetic Minority Over-sampling Technique (SMOTE) were employed to enhance minority class recognition. The proposed Hybrid LSTM-KAN model achieves an overall accuracy of 94.6 percent and a macro-averaged F1 score of 0.703, despite the dominant COPD class accounting for over 86 percent of the data. Improved detection performance is observed for minority classes compared to baseline approaches, demonstrating the effectiveness of the proposed architecture for imbalanced respiratory sound classification.
研究の動機と目的
- 呼吸音データセットにおける深刻なクラス不均衡の課題に対処する。
- 逐次特徴抽出のための LSTM と分類のための Kolmogorov-Arnold Network (KAN) を組み合わせたハイブリッドアーキテクチャを開発する。
- 不均衡緩和技術を特徴抽出と統合して、マイノリティクラスの認識を改善する。
- COPD が支配的な公開六クラス呼吸音データセットで性能を評価する。
提案手法
- LSTM を用いて逐次音声特徴をエンコードする。
- 分類のために LSTM を Kolmogorov-Arnold Network (KAN) と組み合わせる。
- クラス不均衡に対処するため焦点損失を適用する。
- クラス固有のデータ拡張戦略を組み込む。
- マイノリティクラスのバランスをとるため SMOTE を適用する。
- 分布が歪んだ六クラス呼吸音データセットで評価する。
実験結果
リサーチクエスチョン
- RQ1ハイブリッド LSTM-KAN アーキテクチャは、ベースラインと比較して不均衡な呼吸音データの分類を改善できるか。
- RQ2不均衡緩和技術(焦点損失、拡張、SMOTE)はマイノリティクラスの性能にどう影響するか。
- RQ3COPD が支配的な六クラス呼吸音データセットにおける全体の精度とマクロ平均 F1 はどの程度か。
主な発見
- ハイブリッド LSTM-KAN は全体精度 94.6% を達成。
- マクロ平均 F1 スコアは 0.703 に達する。
- マイノリティクラスの検出はベースライン手法と比較して改善した。
- COPD クラスはデータの 86% 超を占めており、深刻な不均衡を浮き彫りにしている。
- 不均衡緩和技術はマイノリティクラスの認識改善に寄与した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。