Skip to main content
QUICK REVIEW

[论文解读] Foundations of Descriptive and Inferential Statistics

Henk van Elst|arXiv (Cornell University)|Feb 11, 2013
Forecasting Techniques and Applications参考文献 61被引用 11
一句话总结

本系列全面的讲义为社会科学、经济学和金融服务业的本科生及研究生提供了技术严谨 yet易于理解的描述统计与推断统计导论。内容涵盖数据描述、概率论、李克特量表、假设检验及线性回归,并整合了R、SPSS、Excel和OpenOffice的软件命令,强调实际应用与效应量解释。

ABSTRACT

These lecture notes were written with the aim to provide an accessible though technically solid introduction to the logic of systematical analyses of statistical data to both undergraduate and postgraduate students, in particular in the Social Sciences, Economics, and the Financial Services. They may also serve as a general reference for the application of quantitative--empirical research methods. In an attempt to encourage the adoption of an interdisciplinary perspective on quantitative problems arising in practice, the notes cover the four broad topics (i) descriptive statistical processing of raw data, (ii) elementary probability theory, (iii) the operationalisation of one-dimensional latent statistical variables according to Likert's widely used scaling approach, and (iv) null hypothesis significance testing within the frequentist approach to probability theory concerning (a) distributional differences of variables between subgroups of a target population, and (b) statistical associations between two variables. The relevance of effect sizes for making inferences is emphasised. These lecture notes are fully hyperlinked, thus providing a direct route to original scientific papers as well as to interesting biographical information. They also list many commands for running statistical functions and data analysis routines in the software packages R, SPSS, EXCEL and OpenOffice. The immediate involvement in actual data analysis practices is strongly recommended.

研究动机与目标

  • 为社会科学、经济学及金融服务业的学生提供坚实且技术可靠的统计方法基础。
  • 通过直接集成R、SPSS、Excel和OpenOffice等软件,弥合理论概念与实际数据分析之间的鸿沟。
  • 通过整合描述统计、概率、测量理论与推断检验,促进定量研究的跨学科方法。
  • 强调效应量与实际意义的重要性,超越p值在假设检验中的作用。
  • 通过嵌入原始研究与人物传记的超链接,并包含可执行代码片段,支持主动学习。

提出的方法

  • 采用结构化、模块化方法,涵盖四个核心领域:单变量数据描述、概率论、潜在变量测量(李克特量表)及零假设显著性检验。
  • 运用频率学派概率方法,用于检验分布差异及变量间的关联性。
  • 整合关键统计技术:单变量频次分布、集中趋势与离散程度度量、偏度、峰度及集中程度(基尼系数)。
  • 采用最小二乘法进行线性回归建模,并推导决定系数(R²)作为模型拟合度量。
  • 通过柯尔莫哥洛夫公理、条件概率及贝叶斯定理,引入形式化概率论。
  • 通过嵌入R、SPSS、Excel和OpenOffice的命令实现实际应用,支持即时动手数据分析。

实验结果

研究问题

  • RQ1如何通过频次分布与汇总统计对原始数据进行系统性描述?
  • RQ2针对不同测量尺度水平,应采用哪些适当的集中趋势、变异性与形状(偏度、峰度)度量?
  • RQ3如何在名义、顺序与尺度量表上量化两变量之间的关联性?
  • RQ4最小二乘法如何生成最佳拟合回归线?R²又如何反映模型性能?
  • RQ5概率论如何为推断决策提供基础,特别是在假设检验与p值解释方面?

主要发现

  • 最小二乘法生成的回归线通过最小化残差平方和,为线性预测提供了稳健基础。
  • 决定系数(R²)量化了因变量中由自变量解释的方差比例,取值范围为0至1。
  • 效应量被强调为p值的重要补充,用于评估实际意义,特别是在组间差异与关联性的假设检验中。
  • 归一化的基尼系数提供了一种尺度不变的集中程度度量,适用于评估分布中的不平等性。
  • 条件概率与贝叶斯定理为基于观测数据更新信念提供了正式框架,支撑了在频率学派语境下的贝叶斯推理。
  • 变量的标准差转换(Z得分)使不同测量尺度间的比较成为可能,有助于多变量分析中的解释。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。