Skip to main content
QUICK REVIEW

[논문 리뷰] Regression with genuinely functional errors-in-covariates

Anirvan Chakraborty, Victor M. Panaretos|arXiv (Cornell University)|2017. 12. 12.
Statistical Methods and Inference참고 문헌 23인용 수 5
한 줄 요약

이 논문은 독립적이고 동일한 분포를 가진 오차가 아닌 일반적인 스토케스틱 과정으로 모델링된 진정한 기능적 측정 오차를 가진 기능적 선형 모형에 대해 새로운 회귀 校정 추정기법을 제안한다. 오차 공분산 구조가 알려져 있지 않은 상황에서 행렬 완성 기법을 활용하여 이를 추정함으로써, 일관된 기울기 추정이 가능해지고, 특히 비i.i.i.d. 오차 구조에서 유한 표본 및 실제 데이터에서 스펙트럼 절단보다 뚜렷이 뛰어난 성능을 발휘한다.

ABSTRACT

Contamination of covariates by measurement error is a classical problem in multivariate regression, where it is well known that failing to account for this contamination can result in substantial bias in the parameter estimators. The nature and degree of this effect on statistical inference is also understood to crucially depend on the specific distributional properties of the measurement error in question. When dealing with functional covariates, measurement error has thus far been modelled as additive white noise over the observation grid. Such a setting implicitly assumes that the error arises purely at the discrete sampling stage, otherwise the model can only be viewed in a weak (stochastic differential equation) sense, white noise not being a second-order process. Departing from this simple distributional setting can have serious consequences for inference, similar to the multivariate case, and current methodology will break down. In this paper, we consider the case when the additive measurement error is allowed to be a valid stochastic process. We propose a novel estimator of the slope parameter in a functional linear model, for scalar as well as functional responses, in the presence of this general measurement error specification. The proposed estimator is inspired by the multivariate regression calibration approach, but hinges on recent advances on matrix completion methods for functional data in order to handle the nontrivial (and unknown) error covariance structure. The asymptotic properties of the proposed estimators are derived. We probe the performance of the proposed estimator of slope using simulations and observe that it substantially improves upon the spectral truncation estimator based on the erroneous observations, i.e., ignoring measurement error. We also investigate the behaviour of the estimators on a real dataset on hip and knee angle curves during a gait cycle.

연구 동기 및 목표

  • 기존 방법들이 기능적 예측 변수에서 i.i.d. 또는 화이트 노이즈 측정 오차를 가정하는 데서 비롯되는 한계를 해결하기 위해.
  • 측정 오차가 일반적인 두 번째 모멘트 스토케스틱 과정인 경우 스칼라-및 기능적-기능 회귀에서 기울기 파rameter에 대한 일관된 추정기법을 개발하기 위해.
  • 기능적 데이터에서 알려져 있지 않은 복잡한 오차 공분산 구조로 인해 발생하는 식별 불가능성 및 불안정성 문제를 다루기 위해.
  • 측정 오차가 i.i.i.d. 가정을 벗어나는 경우 스펙트럼 절단보다 추정 정확도를 향상시키기 위해.
  • 비균일한 측정 오차 분산을 가진 실제 걸음걸이 데이터에서 본 방법의 강건성과 실용적 유용성을 입증하기 위해.

제안 방법

  • 다변량 회귀 校정 프레임워크를 기능적 설정에 적응시켜, 두 단계 접근법을 사용한다: 먼저 진짜 예측 변수 과정을 추정하고, 그 다음 기울기 추정에서 측정 오차를 교정한다.
  • 기능적 측정 오차의 알려지지 않은 공분산 구조를 추정하기 위해 행렬 완성 기법을 활용하여 일반적인 오차 사양 하에서도 일관된 추론이 가능하도록 한다.
  • 진짜 예측 변수 및 오차 과정을 낮은 랭크 부분공간에 표현하기 위해 기능적 주성분 분석(FPCA)을 사용하며, 랭크는 데이터 기반 선택 기준에 의해 결정된다.
  • 두 단계 추정 절차를 구현한다: (1) 스무딩 및 수축을 통해 관측된 데이터에서 오차 공분산을 추정하고, (2) 회귀 校정을 적용하여 편향 보정된 기울기 추정치를 도출한다.
  • 진짜 예측 변수의 효과적 랭크를 선택하기 위해 데이터 기반 스펙트럼 절단 규칙을 적용하여 고차원 설정에서 추정기의 안정성을 확보한다.
  • 이상치에 강건한 페널라이제이션 최대우도 또는 스무딩 접근법을 사용하여 측정 오차 분산 함수를 추정함으로써 오차 과정의 이질성도 수용할 수 있다.

실험 결과

연구 질문

  • RQ1비i.i.i.d. 측정 오차를 가진 기능적 선형 모형에 대해 회귀 校정 접근법을 확장할 수 있는가?
  • RQ2측정 오차가 화이트 노이즈가 아닐 경우, 제안된 추정기의 성능은 스펙트럼 절단과 비교해 어떻게 되는가?
  • RQ3행렬 완성 기법은 측정 오차가 존재하는 기능적 자료에서 알려지지 않은 오차 공분산 구조를 효과적으로 복원할 수 있는가?
  • RQ4식별 불가능성과 고랭크 오차 구조가 기울기 추정에 미치는 영향은 무엇이며, 이를 어떻게 완화할 수 있는가?
  • RQ5제안된 방법은 복잡한 오차 구조를 가진 실제 기능적 자료에서 더 높은 예측 정확도를 제공하는가?

주요 결과

  • 측정 오차가 i.i.i.d. 또는 이종분산이 아닐 경우, 회귀 校정 추정기는 스펙트럼 절단 추정기보다 편향과 평균제곱오차 측면에서 뚜렷이 뛰어나다.
  • 모의 실험에서 무한랭크 설정에서도 진짜 기울기 함수를 성공적으로 복원하였으며, 추정된 필수 랭크가 진짜 기반 구조와 일치함을 확인하였다.
  • 걸음걸이 데이터셋에서, 회귀 校정 추정기는 예측에 대해 R² 54.2%를 기록하였고, 관측된(오차가 있는) 예측 변수를 사용한 스펙트럼 절단 추정기는 50.3%에 그쳤다.
  • 걸음걸이 자료에서 추정된 측정 오차 분산은 강한 이종분산성을 보였으며, 이는 i.i.d. 오차 가정이 성립하지 않음을 시사하고, 본 방법의 필요성을 정당화한다.
  • 고차원 고유함수 구조가 존재하더라도 효과적 랭크 선택 덕분에 추정기는 안정적이고 적응 가능성을 보였다.
  • 오차 구조의 잘못된 특정화에 대해 강건하며, 관측 그리드가 조밀하거나 희박한 경우에도 양호한 성능을 유지한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.