Skip to main content
QUICK REVIEW

[论文解读] Know your population and know your model: Using model-based regression and poststratification to generalize findings beyond the observed sample

Lauren Kennedy, Andrew Gelman|arXiv (Cornell University)|Jun 26, 2019
Mental Health Research Topics被引用 5
一句话总结

本文提出使用多层回归与后分层化(MRP)方法,通过建模处理效应与人口统计变量的交互作用,并利用人口层面的人口统计分布调整预测,将非代表性心理样本的研究发现推广至更广泛人群。该方法通过校正样本偏差,实现了对人格特质分布(如尽责性与神经质)的精确估计,表明MRP能显著提升心理学研究的推广性。

ABSTRACT

Psychology research focuses on interactions, and this has deep implications for inference from non-representative samples. For the goal of estimating average treatment effects, we propose to fit a model allowing treatment to interact with background variables and then average over the distribution of these variables in the population. This can be seen as an extension of multilevel regression and poststratification (MRP), a method used in political science and other areas of survey research, where researchers wish to generalize from a sparse and possibly non-representative sample to the general population. In this paper, we discuss areas where this method can be used in the psychological sciences. We use our method to estimate the norming distribution for the Big Five Personality Scale using open source data. We argue that large open data sources like this and other collaborative data sources can be combined with MRP to help resolve current challenges of generalizability and replication in psychology.

研究动机与目标

  • 解决从非代表性心理样本推广实验发现至更广泛人群的挑战。
  • 证明心理学中的处理效应本质上依赖于人群,因其在不同人口统计子群体中存在异质性。
  • 倡导将MRP——一种来自调查研究的方法——作为心理学研究的标准工具,以提升对便利样本之外的推断能力。
  • 表明开放数据源结合MRP可解决心理学研究中可重复性与推广性的问题。
  • 提供一个实用框架,利用开放源代码数据通过MRP估计人格特质的总体水平分布。

提出的方法

  • 该方法使用多层回归建模处理效应,包含处理条件与人口统计变量(如年龄、性别)之间的交互作用。
  • 利用美国社区调查(ACS)的人口统计数据进行后分层化,将样本估计值重新对齐至真实的人口分布。
  • 拟合一个贝叶斯分层模型,结合弱信息先验,以稳定小样本子群体中的估计值。
  • 为从ACS人口分布中按比例抽取的模拟总体样本生成预测。
  • 使用后验预测分布估计总体中各群体的期望得分,通过多次抽样量化不确定性。
  • 该方法比较样本与总体的预测分布,对人口统计不平衡进行调整。

实验结果

研究问题

  • RQ1心理学研究如何将非代表性样本的发现推广至更广泛人群?
  • RQ2人口统计交互作用在多大程度上影响心理学实验中处理效应的估计?
  • RQ3当应用于开放源代码心理数据时,MRP能否提升人格特质分布估计的准确性?
  • RQ4与原始样本数据相比,MRP调整如何影响大五人格特质分布的估计?
  • RQ5开放数据源与人口基准在提升心理学研究的推广性与可重复性方面发挥什么作用?

主要发现

  • MRP显著调整了尽责性与神经质估计分布,表明原始样本存在显著偏差。
  • 对开放性与宜人性的调整较小,表明其样本分布相对接近总体分布。
  • 外向性在样本与总体估计之间差异可忽略,表明该特质分布偏差极小。
  • 该方法通过后验预测模拟成功生成了带有不确定性量化的总体水平预测。
  • 利用开放数据(如ACS)结合MRP,即使原始样本非代表性,研究人员也能估计总体分布。
  • 结果表明,忽略人口统计异质性会导致平均处理效应估计产生误导,而MRP能有效纠正这一问题。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。