Skip to main content
QUICK REVIEW

[论文解读] How sure are we? Two approaches to statistical inference

Michael Wood|arXiv (Cornell University)|Mar 15, 2018
Diversity and Career in Medicine参考文献 9被引用 5
一句话总结

本文提出了两种统计推断方法——基准假设检验和基于重抽样与自助法的贝叶斯式方法——通过直观的模拟方法评估基于样本结论的不确定性,而无需依赖复杂的概率分布。该方法强调非专业人士的可及性,通过电子表格模拟实现,提供传统零假设检验所缺乏的概率性假设解释。

ABSTRACT

Suppose you are told that taking a statin will reduce your risk of a heart attack or stroke by 3% in the next ten years, or that women have better emotional intelligence than men. You may wonder how accurate the 3% is, or how confident we should be about the assertion about women's emotional intelligence, bearing in mind that these conclusions are only based on samples of data? My aim here is to present two statistical approaches to questions like these. Approach 1 is often called null hypothesis testing but I prefer the phrase "baseline hypothesis": this is the standard approach in many areas of inquiry but is fraught with problems. Approach 2 can be viewed as a generalisation of the idea of confidence intervals, or as the application of Bayes' theorem. Unlike Approach 1, Approach 2 provides a tentative estimate of the probability of hypotheses of interest. For both approaches, I explain, from first principles, building only on "common sense" statistical concepts like averages and randomness, both how to derive answers, and the rationale behind the answers. This is achieved by using computer simulation methods (resampling and bootstrapping using a spreadsheet available on the web) which avoid the use of probability distributions (t, normal, etc). Such a minimalist, but reasonably rigorous, analysis is particularly useful in a discipline like statistics which is widely used by people who are not specialists. My intended audience includes both statisticians, and users of statistical methods who are not statistical experts.

研究动机与目标

  • 为解决在现实世界数据结论中解释统计不确定性的常见挑战,例如风险降低或组间差异。
  • 为非统计学专业人士提供一种比传统零假设检验更直观、更易访问的替代方法。
  • 展示如何通过重抽样与自助法估计假设为真的概率,而无需依赖t分布或正态分布等理论分布。
  • 通过仅使用平均值和随机性等基本概念,弥合统计严谨性与实际理解之间的差距。
  • 提供一个透明、可网络访问的电子表格工具,使用户能够模拟并可视化统计推断过程。

提出的方法

  • 使用计算机模拟(重抽样与自助法)估计抽样分布,而无需假设理论概率分布。
  • 采用“基准假设”框架作为零假设显著性检验的替代方案,聚焦于实际效应大小和不确定性。
  • 使用基于电子表格的模拟环境,使用户能够通过重复随机抽样探索统计推断。
  • 通过在不同假设下模拟数据并观察结果出现的频率,来建立对假设的信心。
  • 通过模拟后验估计方法引入类似贝叶斯的解释,估算在观察到数据的前提下假设为真的概率。
  • 仅使用平均值和随机性等常识性概念,从基本原理出发构建理解,避免复杂的数学形式化。

实验结果

研究问题

  • RQ1如何基于样本数据评估报告的风险降低(如他汀类药物使心脏病发作风险降低3%)的可靠性?
  • RQ2如何恰当地量化对组间比较(如女性情绪智力高于男性)的信心?
  • RQ3非专业人士如何在不依赖高级统计理论的情况下理解并解释统计不确定性?
  • RQ4基准假设方法与传统零假设检验在哪些方面不同?它在哪些方面有所改进?
  • RQ5基于模拟的方法能否提供更直观、更准确的假设为真的概率评估?

主要发现

  • 尽管基准假设方法被广泛使用,但其二元决策框架和对p值的依赖常导致误解。
  • 基于模拟的方法可直接估计假设为真的概率,提供更直观、更具信息量的不确定性度量。
  • 重抽样与自助法使用户能够经验性地推导抽样分布,避免对正态性或t分布的假设。
  • 该方法使用户能够通过重复模拟可视化和理解估计值的变异性,从而增强概念清晰度。
  • 可网络访问的电子表格工具使用户能够交互式地探索统计推断,使复杂概念对非专家更易理解。
  • 该方法成功地将抽象的统计概念转化为实际、可视化且计算基础扎实的见解,适用于实际应用用户。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。