[论文解读] Using a Birth-Death Process to Account for Reporting Errors in Longitudinal Self-reported Counts of Behavior
本文提出了一种贝叶斯分层模型,将线性出生死亡过程与泊松随机效应模型相结合,以校正HIV阳性青年纵向自报性伴侣数量中的报告错误。通过将报告数量建模为未观测到的真实数量的噪声观测值,该方法在忽略回忆误差的标准泊松模型基础上,显著提高了推断准确性,尤其在高频率报告中表现更优。
We analyze longitudinal self-reported counts of sexual partners from youth living with HIV. In self-reported survey data, subjects recall counts of events or behaviors such as the number of sexual partners or the number of drug uses in the past three months. Subjects with small counts may report the exact number, whereas subjects with large counts may have difficulty recalling the exact number. Thus, self-reported counts are noisy, and mis-reporting induces errors in the count variable. As a naive method for analyzing self-reported counts, the Poisson random effects model treats the observed counts as true counts and reporting errors in the outcome variable are ignored. Inferences are therefore based on incorrect information and may lead to conclusions unsupported by the data. We describe a Bayesian model for analyzing longitudinal self-reported count data that formally accounts for reporting error. We model reported counts conditional on underlying true counts using a linear birth-death process and use a Poisson random effects model to model the underlying true counts. A regression version of our model can identify characteristics of subjects with greater or lesser reporting error. We demonstrate several approaches to prior specification.
研究动机与目标
- 解决自报纵向行为计数中的回忆误差问题,特别是高频行为。
- 开发一种统计模型,以区分真实潜在计数与受记忆和报告偏差影响的报告计数。
- 通过正式考虑计数数据中的报告误差,而非假设报告值准确,从而改进公共卫生研究中的推断。
- 通过模型的回归扩展,识别与更高或更低报告误差相关的个体水平预测因子。
- 提供一个具有恰当先验设定和MCMC抽样方法的灵活贝叶斯框架,用于后验推断。
提出的方法
- 将观测到的自报计数 $Y_{ij}$ 建模为给定未观测真实计数 $Z_{ij}$ 的条件分布,使用具有速率参数 $\lambda_{ij}$ 的线性出生死亡(BD)过程。
- 使用泊松随机效应模型(PREM)对真实计数 $Z_{ij}$ 建模,纳入个体特定的随机截距 $\epsilon_i$ 和协变量。
- 采用分层贝叶斯框架,其中 $\lambda_{ij} = \exp(\mathbf{w}_{ij}^\top \boldsymbol{\psi} + \epsilon_i)$,通过在对数率上进行回归,将BD过程与协变量关联。
- 应用Metropolis-Hastings MCMC抽样方法估计潜在真实计数 $Z_{ij}$,并使用状态相关提议分布以确保有效整数状态转移。
- 对变异分量使用共轭先验,例如对 $D_\epsilon^{-1}$ 使用逆伽马先验,对回归系数使用扩散或弱信息先验。
- 通过将BD似然与PREM似然结合,推导联合后验密度,并对较大的 $Z_{ij}$ 值引入对数伽马校正。
实验结果
研究问题
- RQ1在自报性伴侣数量中考虑报告误差,如何影响纵向行为研究中的推断?
- RQ2报告误差在多大程度上随真实计数的大小而变化?是否可以对其进行随机建模?
- RQ3出生死亡过程能否有效捕捉自报行为计数中回忆误差的随机动态?
- RQ4哪些个体水平特征与更高或更低的报告误差相关?如何对其进行估计?
- RQ5与标准泊松随机效应模型相比,所提出的模型在偏差和精度方面表现如何?
主要发现
- 该模型通过将报告计数视为未观测真实计数的噪声观测值,成功校正了报告误差,降低了参数估计的偏差。
- 出生死亡过程为建模高计数下过度或不足报告的概率提供了灵活的随机机制。
- 在真实计数模型中引入个体特定的随机效应,可捕捉未观测到的异质性及时间点间的相关性。
- 该模型的回归扩展识别出与报告误差相关的协变量,从而可识别出高误差亚群。
- 通过使用状态相关提议的Metropolis-Hastings后验抽样,确保了对整数值潜在计数空间的有效探索。
- 该模型通过避免基于错误且易出错的数据进行推断,优于朴素的泊松模型。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。