[论文解读] Random weighted averages, partition structures and generalized arcsine laws
本文提出了一种基于分划结构和广义反正弦定律的随机加权平均(P-均值)分布理论的统一框架。它建立了将P-均值与从P中抽取样本的相异值数量的概率生成函数联系起来的矩公式,并推导出(α,θ)-均值的显式柯西-史蒂尔杰斯变换,其中包含广义反正弦定律和狄利克雷均值作为特例。
This article offers a simplified approach to the distribution theory of randomly weighted averages or $P$-means $M_P(X):= \sum_{j} X_j P_j$, for a sequence of i.i.d.random variables $X, X_1, X_2, \ldots$, and independent random weights $P:= (P_j)$ with $P_j \ge 0$ and $\sum_{j} P_j = 1$. The collection of distributions of $M_P(X)$, indexed by distributions of $X$, is shown to encode Kingman's partition structure derived from $P$. For instance, if $X_p$ has Bernoulli$(p)$ distribution on $\{0,1\}$, the $n$th moment of $M_P(X_p)$ is a polynomial function of $p$ which equals the probability generating function of the number $K_n$ of distinct values in a sample of size $n$ from $P$: $E (M_P(X_p))^n = E p^{K_n}$. This elementary identity illustrates a general moment formula for $P$-means in terms of the partition structure associated with random samples from $P$, first developed by Diaconis and Kemperman (1996) and Kerov (1998) in terms of random permutations. As shown by Tsilevich (1997) if the partition probabilities factorize in a way characteristic of the generalized Ewens sampling formula with two parameters $(α,θ)$, found by Pitman (1992), then the moment formula yields the Cauchy-Stieltjes transform of an $(α,θ)$ mean. The analysis of these random means includes the characterization of $(0,θ)$-means, known as Dirichlet means, due to Von Neumann (1941), Watson (1956) and Cifarelli and Regazzini (1990) and generalizations of Lévy's arcsine law for the time spent positive by a Brownian motion, due to Darling (1949) Lamperti (1958) and Barlow, Pitman and Yor (1989).
研究动机与目标
- 开发P-均值分布的简化基础理论,P-均值定义为加权平均∑XjPj,其中Pj≥0且总和为1。
- 阐明P-均值与金曼分划结构之间的联系,表明P-均值的矩编码了从P中抽取随机样本时相异值的分布。
- 通过柯西-史蒂尔杰斯变换刻画(α,θ)-均值的分布,扩展经典关于反正弦定律和狄利克雷均值的结果。
- 在基于矩的统一框架下,整合此前分散的结果,如莱维的反正弦定律和冯·诺伊曼的狄利克雷均值。
- 提供易于理解的推导过程,并与文献建立联系,尤其面向不熟悉分划结构及其在P-均值理论中作用的读者。
提出的方法
- 使用矩公式E[M_P(X)]^n = E[p^{K_n}],其中X为伯努利(p)变量,将P-均值的矩与从P中抽取大小为n的样本时相异值数量K_n的概率生成函数联系起来。
- 应用柯西-史蒂尔杰斯变换来刻画(α,θ)-均值的分布,推导出关键恒等式E[(1+λX̃)^{-θ}] = (E[(1+λX)^α])^{-θ/α},其中α≠0,θ≠0。
- 利用单调收敛定理并借助简单随机变量的逼近,将结果从有界变量推广至非负无界X≥0的情形。
- 利用GEM(α,θ)对P的表示以及两参数泊松-狄利克雷过程,推导出广义反正弦定律作为极限情形。
- 应用碎片化算子和中国餐馆过程中的反向鞅理论,分析P-均值的结构。
- 通过形式幂级数展开推导显式矩公式,通过比较生成函数中λ^n的系数来计算E[X̃_α,θ^n]。
实验结果
研究问题
- RQ1如何通过从随机权重P中抽样所诱导的分划结构来刻画P-均值的分布?
- RQ2P-均值的矩与从P中抽取样本时相异值数量的生成函数之间的确切关系是什么?
- RQ3(α,θ)-均值分布的柯西-史蒂尔杰斯变换是什么?它如何统一已知结果,如狄利克雷均值和广义反正弦定律?
- RQ4在何种X和(α,θ)条件下,(α,θ)-均值几乎必然有限或无限?
- RQ5结果如何从有界非负随机变量X推广至无界非负随机变量?
主要发现
- 对于伯努利(p)的X,M_P(X_p)的n阶矩为E[p^{K_n}],其中K_n为从P中抽取大小为n的样本时的相异值数量,建立了P-均值矩与分划结构之间的直接联系。
- 对于0<α<1且θ>−α的(α,θ)-均值,若E[X^α]<∞,则X̃_α,θ几乎必然有限;若E[X^α]=∞,则几乎必然无限,给出了精确的可积性条件。
- 当α≠0,θ≠0时,(α,θ)-均值的柯西-史蒂尔杰斯变换为E[(1+λX̃_α,θ)^{-θ}] = (E[(1+λX)^α])^{-θ/α},该式推广了狄利克雷分布和反正弦定律的已知变换。
- 当θ=0时,得到E[log(1+λX̃_α,0)] = (1/α)log(E[(1+λX)^α]),当X为指示变量时,该式退化为广义反正弦定律的兰珀蒂变换。
- 公式E[(1+λX̃_α,0)^{-1}] = E[(1+λX)^{α-1}]/E[(1+λX)^α]提供了变换的另一种形式,适用于0<α<1,并在X简单时恢复巴洛等人的结果。
- 该理论将对称狄利克雷均值变换(141)和无限狄利克雷均值变换(144)统一为一般(α,θ)-均值公式的特例,后者在α→0+时出现。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。