Skip to main content
QUICK REVIEW

[论文解读] Approximation Theorems Related to the Coupon Collector's Problem

Anna Pósfai|arXiv (Cornell University)|Jun 17, 2010
Markov Chains and Monte Carlo Methods参考文献 26被引用 7
一句话总结

本篇哲学博士论文利用Stein方法与耦合技术,对优惠券收集者等待时间 $W_{n,m}$ 建立了高精度的近似定理,以改进经典极限定理。论文在中心极限与泊松极限情形下,建立了误差界为 $1/n$ 阶的泊松、复合泊松及泊松–夏勒尔展开,显著优于以往的 $1/\sqrt{n}$ 速率,尤其在这些极限区域表现更优。

ABSTRACT

This Ph.D. thesis concerns the version of the classical coupon collector's problem, when a collector samples with replacement a set of $n\ge 2$ distinct coupons so that at each time any one of the $n$ coupons is drawn with the same probability $1/n$. For a fixed integer $m\in\{0,1,...,n-1\}$, the coupon collector's waiting time $W_{n,m}$ is the random number of draws the collector performs until he acquires $n-m$ distinct coupons for the first time. The basic goal of the thesis is to approximate the distribution of the coupon collector's appropriately centered and normalized waiting time with well-known measures with high accuracy, and in many cases prove asymptotic expansions for the related probability distribution functions and mass functions. The approximating measures are chosen from five different measure families. Three of them -- the Poisson distributions, the normal distributions and the Gumbel-like distributions -- are probability measure families whose members occur as limiting laws in the limit theorems concerning $W_{n,m}$. The other two approximating measure families are certain compound Poisson distributions and Poisson--Charlier signed measures.

研究动机与目标

  • 通过实现更高阶的近似精度,改进优惠券收集者等待时间 $W_{n,m}$ 的经典极限定理。
  • 弥补现有近似方法的不足,提供优于标准 $1/\sqrt{n}$ 速率的更紧误差界。
  • 发展并应用先进的耦合与Stein方法技术,以在中心极限与泊松极限情形下近似分布。
  • 为中心化等待时间 $\widetilde{W}_{n,m} = W_{n,m} - (n-m)$ 建立显式误差率的泊松–夏勒尔展开。
  • 比较复合泊松与正态近似的性能,表明在某些参数区域中,复合泊松近似可与正态近似同样或更加精确。

提出的方法

  • 应用Stein方法,以有界总变差距离来控制 $W_{n,m}$ 的分布与泊松、正态及复合泊松近似之间的差异。
  • 提出一种新颖的耦合不等式,以在各项方差随 $n$ 递减的情况下改进误差界。
  • 利用特征函数技术,推导中心化等待时间 $\widetilde{W}_{n,m}$ 的泊松–夏勒尔展开。
  • 根据 $m/n$ 与 $\sigma_n^2 - \mu_n - (n-m)$ 的渐近行为,建立误差界为 $\sigma_n \log \sigma_n (1/\sqrt{m})^R$、$(\sqrt{n})^{R-3}/(n-m)^{R-2}$ 与 $(n-m)/(√n)^{R+1}$ 的误差界。
  • 利用独立几何随机变量的表示式 $W_{n,m} \stackrel{d}{=} \sum_{k=m+1}^n X_{k/n}$ 来分析各阶矩与依赖结构。
  • 以恒等式 $\mathbb{E}[W_{n,m}] = n \sum_{k=m+1}^n \frac{1}{k}$ 与 $\mathrm{Var}(W_{n,m}) = n \sum_{k=m+1}^{n-1} \frac{n-k}{k^2}$ 作为归一化与中心化的基础工具。

实验结果

研究问题

  • RQ1能否使 $\widetilde{W}_{n,m} = W_{n,m} - (n-m)$ 的泊松近似达到 $O(1/n)$ 的误差率,从而优于经典的 $O(1/\sqrt{n})$ 界?
  • RQ2在中心极限与泊松极限情形下,复合泊松近似与正态近似在总变差距离上的误差表现如何比较?
  • RQ3中心化等待时间 $\widetilde{W}_{n,m}$ 的泊松–夏勒尔展开的最优收敛速率是多少?其依赖于 $m/n$ 的渐近行为如何?
  • RQ4是否可通过新提出的耦合不等式,在各项方差随 $n$ 递减的情形(如优惠券收集过程)中改进误差界?
  • RQ5在何种条件下,复合泊松近似可实现优于 $O(1/\sqrt{\mathrm{Var}(W_{n,m})})$ 的误差率?

主要发现

  • 对 $\widetilde{W}_{n,m} = W_{n,m} - (n-m)$ 的泊松近似在总变差距离下实现了 $1/n$ 阶的误差界,显著优于经典的 $1/\sqrt{n}$ 速率。
  • 复合泊松近似对 $W_{n,m}$ 的总变差误差界为 $O(1/\sqrt{n})$,但通过新提出的耦合不等式,显著改善了常数项,尤其在各项方差随 $n$ 递减时表现更优。
  • 对于泊松–夏勒尔展开,误差率取决于 $m/n$ 与 $\sigma_n^2 - \mu_n - (n-m)$ 的渐近行为,分别为 $O(\sigma_n \log \sigma_n (1/\sqrt{m})^R)$、$O((\sqrt{n})^{R-3}/(n-m)^{R-2})$ 或 $O((n-m)/(√n)^{R+1})$,其中 $R \geq 3$。
  • 复合泊松近似在中心极限与泊松极限情形下,其总变差距离误差与正态近似相当或更优。
  • 结合均值匹配的Stein方法与新提出的耦合不等式,可获得比经典Mineka不等式更紧的界,尤其在非i.i.d.情形且方差递减时表现更优。
  • 通过推导极限分布的精确密度函数,对类似Gumbel的近似进行了改进,其位置参数偏移为 $\gamma - \sum_{k=1}^m \frac{1}{k}$,且当 $m$ 固定时该近似成立。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。