[论文解读] Statistical Inference for Rényi Entropy Functionals
本文提出了一种基于U-统计量的新型估计器,用于在独立同分布样本中,利用ε-邻近向量记录,对离散和连续分布中的Rényi熵泛函进行估计。在一般条件下,该估计器具有相合性和渐近正态性,从而支持在近似匹配和图像分析等应用中对熵型泛函进行有效的统计推断。
Numerous entropy-type characteristics (functionals) generalizing Rényi entropy are widely used in mathematical statistics, physics, information theory, and signal processing for characterizing uncertainty in probability distributions and distribution identification problems. We consider estimators of some entropy (integral) functionals for discrete and continuous distributions based on the number of epsilon-close vector records in the corresponding independent and identically distributed samples from two distributions. The estimators form a triangular scheme of generalized U-statistics. We show the asymptotic properties of these estimators (e.g., consistency and asymptotic normality). The results can be applied in various problems in computer science and mathematical statistics (e.g., approximate matching for random databases, record linkage, image matching).
研究动机与目标
- 开发一个统一的框架,用于估计离散和连续分布中的Rényi熵泛函 $ q_{\mathbf{r}} = \int p_X^{r_1} p_Y^{r_2} dx $。
- 建立基于i.i.d.样本中ε-邻近记录的核型估计器的渐近性质——相合性和渐近正态性。
- 将先前关于二次Rényi熵估计的工作扩展到更广泛的具有任意指数 $ r_1, r_2 \geq 0 $ 的熵型泛函类。
- 通过渐近置信区间,支持在近似匹配、记录链接和图像匹配等应用中的统计推断。
提出的方法
- 利用来自 $ \mathcal{P}_X $ 和 $ \mathcal{P}_Y $ 的样本中 $ \epsilon $-邻近观测的指示函数,构建广义U-统计量 $ Q_{\mathbf{n}} $。
- 将对称化核 $ \psi_{\mathbf{n}}(S;T) $ 定义为所有 $ r_1 $-子集 $ S $ 的 $ X $-样本与所有 $ r_2 $-子集 $ T $ 的 $ Y $-样本的指示函数的平均值。
- 推导出 $ \epsilon $-匹配概率 $ q_{\mathbf{r},\epsilon} = \mathbb{E}[p_{X,\epsilon}^{r_1-1} p_{Y,\epsilon}^{r_2}] $,即核的期望。
- 应用Lindeberg-Feller中心极限定理,证明归一化估计器 $ n^{1/2}(\tilde{Q}_{\mathbf{n}} - \tilde{q}_{\mathbf{r},\epsilon}) $ 的渐近正态性。
- 通过涉及 $ D_\epsilon(x) $ 的偏差分解,表达 $ \tilde{q}_{\mathbf{r},\epsilon} $ 与 $ q_{\mathbf{r}} $ 之间的差异,以分析收敛速率。
- 在密度满足Hölder连续性假设下建立收敛速率,其中 $ \epsilon $ 作为样本量 $ n $ 的函数进行选择。
实验结果
研究问题
- RQ1如何利用ε-邻近记录统计量,从i.i.d.样本中一致估计Rényi熵泛函?
- RQ2所提出的U-统计量估计器在 $ q_{\mathbf{r}} $ 上的渐近分布性质是什么?
- RQ3当底层密度满足Hölder连续性时,偏差和方差的收敛速率可以达到多快?
- RQ4ε 的选择如何影响 $ q_{\mathbf{r}} $ 估计中的偏差-方差权衡?
- RQ5该估计器在实践中能否用于构建熵泛函的渐近有效置信区间?
主要发现
- 在密度的弱正则性条件下,所提出的估计器 $ \tilde{Q}_{\mathbf{n}} $ 对 $ q_{\mathbf{r}} $ 具有相合性。
- 归一化估计器 $ n^{1/2}(\tilde{Q}_{\mathbf{n}} - \tilde{q}_{\mathbf{r},\epsilon}) $ 依分布收敛于均值为零、方差为 $ \kappa $ 的正态随机变量,从而确立了渐近正态性。
- 当 $ \alpha < d/2 $ 时,偏差与方差的联合阶为 $ \mathrm{O}(n^{-2\alpha/(2\alpha + d(1 - 1/r))}) $,其中 $ \alpha $ 为密度的Hölder指数。
- 当 $ \epsilon \sim c n^{-1/(2\alpha + d(1 - 1/r))} $ 时,估计器在Hölder光滑性下达到最优收敛速率。
- 当 $ \alpha = d/2 $ 时,收敛速率仍为 $ \mathrm{O}(n^{-\alpha/(2\alpha + d(1 - 1/r))}) $,且估计器保持相合性。
- 当 $ \alpha > d/2 $ 时,若 $ \epsilon \sim L(n) n^{-1/d} $ 且 $ n\epsilon^d \to \infty $,则估计器保持相合性和渐近正态性,利用Slutsky定理。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。