Skip to main content
QUICK REVIEW

[論文レビュー] Quantitative model validation techniques: new insights

You Ling, Sankaran Mahadevan|arXiv (Cornell University)|Jun 21, 2012
Probabilistic and Robust Engineering Design参考文献 48被引用数 6
ひとこと要約

本稿では、計算モデルの予測を評価するための4つの定量的モデル妥当性評価手法——古典的およびベイジアン仮説検定、信頼性に基づく手法、および領域メトリックに基づく手法——を提案および評価する。方向性バイアスを考慮するためのベイジアン区間仮説検定を導入し、閾値最適化とモデル平均化を通じてベイジアン手法がタイプI/IIエラーのリスクを低減することを示している。特定の条件下では、ベイジアン手法は古典的p値と強い数学的関係を示す。

ABSTRACT

This paper develops new insights into quantitative methods for the validation of computational model prediction. Four types of methods are investigated, namely classical and Bayesian hypothesis testing, a reliability-based method, and an area metric-based method. Traditional Bayesian hypothesis testing is extended based on interval hypotheses on distribution parameters and equality hypotheses on probability distributions, in order to validate models with deterministic/stochastic output for given inputs. Two types of validation experiments are considered - fully characterized (all the model/experimental inputs are measured and reported as point values) and partially characterized (some of the model/experimental inputs are not measured or are reported as intervals). Bayesian hypothesis testing can minimize the risk in model selection by properly choosing the model acceptance threshold, and its results can be used in model averaging to avoid Type I/II errors. It is shown that Bayesian interval hypothesis testing, the reliability-based method, and the area metric-based method can account for the existence of directional bias, where the mean predictions of a numerical model may be consistently below or above the corresponding experimental observations. It is also found that under some specific conditions, the Bayes factor metric in Bayesian equality hypothesis testing and the reliability-based metric can both be mathematically related to the p-value metric in classical hypothesis testing. Numerical studies are conducted to apply the above validation methods to gas damping prediction for radio frequency (RF) microelectromechanical system (MEMS) switches. The model of interest is a general polynomial chaos (gPC) surrogate model constructed based on expensive runs of a physics-based simulation model, and validation data are collected from fully characterized experiments.

研究の動機と目的

  • 完全に特徴付けられた実験と部分的に特徴付けられた実験、決定論的予測と確率的予測、方向性バイアス、閾値選択といった未解決の課題に対処すること。
  • ベイジアン仮説検定を、分布パラメータの区間仮説および確率密度関数(PDF)の等価性仮説へと拡張し、不確実性下でのモデル妥当性を向上させること。
  • 古典的仮説検定、ベイジアン仮説検定、信頼性に基づく手法、領域メトリックに基づく手法の4つの手法の性能を、異なる実験データ条件下で評価すること。
  • リスクに基づく閾値とモデル平均化を用いることで、ベイジアン手法がモデル選択誤差を最小化することを示すこと。
  • 方向性バイアス(モデル予測が一貫して観測値を低くまたは高く予測する)が、ベイジアン区間検定、信頼性に基づく手法、領域メトリック手法によって定量的に捉えられることを示すこと。

提案手法

  • 予測された平均値と標準偏差を実験データと比較することで、尤度関数と後確率を用いて、予測の妥当性をベイジアン区間仮説検定により評価する手法を提案する。
  • ベイジアン仮説検定を、モデル予測と実験観測の確率密度関数(PDF)の等価性仮説へと拡張する。
  • 許容区間を用いて故障確率を計算し、古典的仮説検定の指標と比較する信頼性に基づく手法を導入する。
  • 予測残差の経験的分布関数(CDF)と一様分布関数(CDF)の距離を計算することで、モデル整合性を定量化する領域メトリックに基づく手法を開発する。
  • 140個の完全に特徴付けられた実験データ点を用いて、RF MEMSスイッチの減衰をモデル化する一般化されたポリノミアルクラウド(gPC)補間モデルにこれらの手法を適用する。
  • ベイジアン妥当性評価において、後確率に基づくモデル平均化を用いることで、タイプI/IIエラーのリスクを低減し、モデル選択のロバストネスを向上させる。

実験結果

リサーチクエスチョン

  • RQ1ベイジアン仮説検定を、モデルパラメータの区間仮説およびPDFの等価性仮説へとどのように拡張できるか。これにより、不確実性下でのモデル妥当性がどのように向上するか。
  • RQ2古典的仮説検定とベイジアン仮説検定は、モデル予測における方向性バイアスに対してどのように感度が異なるか。
  • RQ3信頼性に基づく手法と領域メトリックに基づく手法は、モデル予測における方向性バイアスをどのように検出し、定量的に評価するか。
  • RQ4ベイズ因子が、特定の条件下で古典的p値と数学的に関連づけられる条件は何か。
  • RQ5ベイジアンモデル妥当性の結果は、エンジニアリングシステムの長期的信頼性と故障解析に直接統合可能か。

主な発見

  • ベイジアン区間仮説検定は、モデル予測が一貫して実験データを低くまたは高めに予測する方向性バイアスを効果的に捉え、このような状況では古典的手法よりも性能が劣ることが分かった。
  • 信頼性に基づく手法と領域メトリックに基づく手法の両方のメトリックが、方向性バイアスを反映しており、バイアスがある状況では高い値が示され、モデル性能が悪化していることを示している。
  • RF MEMSスイッチのgPC補間モデルでは、中間圧力(28,664.31 Paおよび43,596.41 Pa)で実験データとの整合性が最も高かったが、最低圧(18,798.45 Pa)および最高圧(66,661.19 Pa)では性能が低下した。
  • 領域メトリックは、18,798.45 PaにおけるgPCモデルを最悪のパフォーマンスを示すと特定し(領域メトリック = 0.543)、主に方向性バイアスに起因していた。
  • z検定と信頼性に基づく手法は、z検定で用いられる有意水準αに基づいて信頼性閾値を設定した場合、同一の故障率を示した。
  • 特定の条件下で、ベイジアン仮説検定のベイズ因子と信頼性に基づくメトリックは、p値と数学的に関連づけられ、古典的およびベイジアン妥当性フレームワークの橋渡しを確立した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。