Skip to main content
QUICK REVIEW

[論文レビュー] Regression with genuinely functional errors-in-covariates

Anirvan Chakraborty, Victor M. Panaretos|arXiv (Cornell University)|Dec 12, 2017
Statistical Methods and Inference参考文献 23被引用数 5
ひとこと要約

本稿は、説明変数に真正の関数的測定誤差(i.i.d. ホワイトノイズではなく一般の確率過程としてモデル化される)を有する関数的線形モデルに対して、新しい回帰補正推定量を提案する。行列補完技術を用いて未知の誤差共分散構造を推定することで、一貫性のある回帰係数推定が可能となり、特に非i.i.d.誤差構造下では、有限標本および実データにおいてスペクトル切断法を著しく上回る性能を示す。

ABSTRACT

Contamination of covariates by measurement error is a classical problem in multivariate regression, where it is well known that failing to account for this contamination can result in substantial bias in the parameter estimators. The nature and degree of this effect on statistical inference is also understood to crucially depend on the specific distributional properties of the measurement error in question. When dealing with functional covariates, measurement error has thus far been modelled as additive white noise over the observation grid. Such a setting implicitly assumes that the error arises purely at the discrete sampling stage, otherwise the model can only be viewed in a weak (stochastic differential equation) sense, white noise not being a second-order process. Departing from this simple distributional setting can have serious consequences for inference, similar to the multivariate case, and current methodology will break down. In this paper, we consider the case when the additive measurement error is allowed to be a valid stochastic process. We propose a novel estimator of the slope parameter in a functional linear model, for scalar as well as functional responses, in the presence of this general measurement error specification. The proposed estimator is inspired by the multivariate regression calibration approach, but hinges on recent advances on matrix completion methods for functional data in order to handle the nontrivial (and unknown) error covariance structure. The asymptotic properties of the proposed estimators are derived. We probe the performance of the proposed estimator of slope using simulations and observe that it substantially improves upon the spectral truncation estimator based on the erroneous observations, i.e., ignoring measurement error. We also investigate the behaviour of the estimators on a real dataset on hip and knee angle curves during a gait cycle.

研究の動機と目的

  • 関数的予測変数における測定誤差がi.i.d.またはホワイトノイズであると仮定する従来の手法の限界を是正すること。
  • 測定誤差が一般の2次モーメントの確率過程である場合に、スカラーおよび関数的予測変数の回帰における回帰係数パラメータの一貫性のある推定量を開発すること。
  • 関数データにおける未知で複雑な誤差共分散構造に起因する識別不能性および不安定性の問題に対処すること。
  • 誤差構造がi.i.d.仮定から逸脱する場合、特にスペクトル切断法よりも推定精度を向上させること。
  • 非一様な測定誤差分散を示す実際の歩行データに対して、本手法の頑健性および実用的有効性を示すこと。

提案手法

  • 多変量回帰補正フレームワークを関数的設定に適応し、2段階のアプローチを採用:まず真の予測変数過程を推定し、次に回帰係数推定における測定誤差を補正する。
  • 行列補完技術を用いて、関数的測定誤差の未知の共分散構造を推定し、一般の誤差仕様下でも一貫性のある推論を可能にする。
  • 関数的主成分分析(FPCA)を用いて、真の予測変数および誤差過程を低ランク部分空間に表現し、ランクはデータ駆動型選択により決定する。
  • 2段階推定手順を実装:(1) スムージングと縮小を用いて観測データから誤差共分散を推定、(2) 回帰補正を適用してバイアス補正済みの回帰係数推定値を取得する。
  • 真の予測変数の有効ランクを特定するデータ駆動型スペクトルカットオフルールを適用し、高次元設定下での推定量の安定化を図る。
  • 測定誤差の分散関数を推定するためのペナルティ付き尤度またはスムージングアプローチを用い、誤差過程における異分散性を許容する。

実験結果

リサーチクエスチョン

  • RQ1非i.i.d.測定誤差を有する関数的線形モデルに、回帰補正アプローチを拡張可能か?
  • RQ2測定誤差がホワイトノイズでない場合、提案手法の推定量はスペクトル切断法と比べてどのように性能を発揮するか?
  • RQ3行列補完法は、測定誤差を伴う関数データにおける未知の誤差共分散構造を効果的に回復できるか?
  • RQ4識別不能性および高ランク誤差構造が回帰係数推定に与える影響は何か? また、その影響をどのように軽減できるか?
  • RQ5提案手法は、複雑な誤差構造を有する実関数データにおいて、予測精度を向上させるか?

主な発見

  • 測定誤差がi.i.d.または均一分散でない場合、回帰補正推定量はバイアスおよび平均二乗誤差の観点で、スペクトル切断推定量を著しく上回る。
  • 無限ランク設定下でも、真の回帰係数関数を効果的に回復でき、シミュレーションでは推定された本質的ランクが真の構造と一致した。
  • 歩行データセットでは、回帰補正推定量が予測のためのR²を54.2%に達したのに対し、観測済み(誤った)予測変数を用いたスペクトル切断推定量は50.3%にとどまった。
  • 歩行データにおける推定された測定誤差分散は強い異分散性を示しており、i.i.d.誤差仮定が不適切であることを裏付け、本手法の必要性を正当化した。
  • 高次の固有関数が真の回帰係数に存在する場合でも、効果的なランク選択のおかげで推定量は安定性と適応性を示した。
  • 誤差構造の誤指定に対しても頑健であり、観測グリッドが密または疎であっても良好な性能を維持した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。