Skip to main content
QUICK REVIEW

[论文解读] Inference on a Distribution Function from Ranked Set Samples

Lutz Duembgen, Ehsan Zamanzade|arXiv (Cornell University)|Apr 25, 2013
Statistical Distribution Estimation and Applications参考文献 13被引用 4
一句话总结

本文针对具有固定或随机排名的排序抽样数据,开发并比较了三种累积分布函数(CDF)估计量——分层估计量、非参数最大似然估计量和矩估计量。研究建立了这三种估计量的函数中心极限定理,表明矩估计量在效率和对不完美排名的鲁棒性方面表现更优,尤其在非平衡抽样设计下优势显著。

ABSTRACT

Consider independent observations $(X_1,R_1)$, $(X_2,R_2)$, \ldots, $(X_n,R_n)$ with random or fixed ranks $R_i \in \{1,2,\ldots,k\}$, while conditional on $R_i = r$, the random variable $X_i$ has the same distribution as the $r$-th order statistic within a random sample of size $k$ from an unknown continuous distribution function $F$. Such observation schemes are utilized in situations in which ranking observations is much easier than obtaining their precise values. Two well-known special cases are ranked set sampling (McIntyre 1952) and judgement post-stratification (MacEachern et al. 2004). Within a general setting including unbalanced ranked set sampling we derive and compare the asymptotic distributions of three different estimators of the distribution function $F$ as $n o \infty$ with fixed $k$: The stratified estimator of Stokes and Sager (1988), the nonparametric maximum-likelihood estimator of Kvam and Samaniego (1994) and a moment-based estimator of Chen (2001). Our functional central limit theorems generalize and refine previous asymptotic analyses. In addition we discuss briefly pointwise and simultaneous confidence intervals for the distribution function $F$ with guaranteed coverage probability for finite sample sizes. The methods are illustrated with a real data example, and the potential impact of imperfect rankings is investigated in a small simulation experiment. All in all, the moment-based estimator seems to offer a good compromise between efficiency and robustness versus imperfect ranking, in addition to computational efficiency.

研究动机与目标

  • 开发并比较来自具有固定或随机排名的排序抽样数据的累积分布函数(CDF)的渐近推断方法。
  • 将现有渐近理论从平衡排序抽样扩展至一般非平衡设计和随机排名方案。
  • 在不完美排名和有限样本条件下,评估三种CDF估计量的相对效率与鲁棒性。
  • 构建具有保证有限样本覆盖概率的点态和同时置信区间。
  • 基于效率、鲁棒性与计算可行性,为估计量选择提供实际指导。

提出的方法

  • 在一般非平衡排序抽样下,推导分层估计量(Stokes和Sager,1988)的渐近分布。
  • 通过条件似然最大化方法分析非参数最大似然估计量(Kvam和Samaniego,1994),其基于贝塔分布的顺序统计量。
  • 提出并研究一种基于矩的估计量(Chen,2001),利用估计方程匹配经验概率与理论概率。
  • 为所有三种估计量建立函数中心极限定理,表明其弱收敛于以分布函数为索引的高斯过程。
  • 利用经验过程理论和贝塔分布性质,推导每种估计量的渐近方差函数。
  • 通过模拟和理论界,利用极限高斯过程构造置信带,并保证有限样本覆盖概率。

实验结果

研究问题

  • RQ1在非平衡排序抽样中,分层估计量、MLE估计量与矩估计量的渐近分布如何比较?
  • RQ2在固定 $k$ 且 $n \to \infty$ 条件下,三种估计量的相对渐近效率如何?
  • RQ3不完美排名在有限样本中如何影响每种估计量的表现?
  • RQ4能否构建具有保证有限样本覆盖概率的CDF同时置信带?
  • RQ5哪种估计量在效率、对排名错误的鲁棒性与计算简便性之间提供了最佳权衡?

主要发现

  • 矩估计量在三种估计量中相对渐近效率最高,尤其在非平衡抽样和不完美排名条件下表现更优。
  • 分层估计量具有最大的渐近方差,当某些层大小 $N_{nr}$ 为零时,其性能显著下降。
  • 非参数MLE估计量比分层估计量对空层更具鲁棒性,但效率低于矩估计量。
  • 小样本模拟实验表明,矩估计量在不完美排名下仍能保持良好性能。
  • 理论与模拟结果共同证实,矩估计量在效率与鲁棒性之间提供了强有力的折中方案。
  • 可通过本文推导的极限高斯过程,构建具有保证有限样本覆盖概率的CDF置信带。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。