Skip to main content
QUICK REVIEW

[论文解读] Combining individually valid and conditionally i.i.d. P-variables

Lutz Mattner|arXiv (Cornell University)|Aug 30, 2010
Algorithms and Data Compression参考文献 4被引用 5
一句话总结

该论文提出了一种方法,通过使用第k阶统计量 $U_{k:n}$ 将给定数据条件下条件独立同分布的个体有效P-变量组合起来,生成一个有效的汇总P-变量。关键结果是,在假设条件下,$V = 1 \land \left(\frac{n}{k} U_{k:n}\right)$ 是一个有效的P-变量,且 $f_{n,k}(u) \leq 1 \land \left(\frac{n}{k}u\right)$ 是满足此性质的最小递增函数,从而在弱依赖假设下提供了一种稳健、分布自由的多重检验方法。

ABSTRACT

For a given testing problem, let $U_1,...,U_n$ be individually valid and conditionally on the data i.i.d.\ P-variables (often called P-values). For example, the data could come in groups, and each $U_i$ could be based on subsampling just one datum from each group in order to satisfy an independence assumption under the hypothesis. The problem is then to deterministically combine the $U_i$ into a valid summary P-variable. Restricting here our attention to functions of a given order statistic $U_{k:n}$ of the $U_i$, we compute the function $f_{n,k}$ which is smallest among all increasing functions $f$ such that $f(U_{k:n})$ is always a valid P-variable under the stated assumptions. Since $f_{n,k}(u)\le 1\wedge (\frac {n}{k} u)$, with the right hand side being a good approximation for the left when $k$ is large, one may in particular always take the minimum of 1 and twice the left sample median of the given P-variables. We sketch the original application of the above in a recent study of associations between various primate species by Astaras et al.

研究动机与目标

  • 解决在观测值在组内存在依赖关系时,将个体有效但仅在给定数据条件下条件独立同分布的多个P-变量进行组合的挑战。
  • 推导一个确定性且置换不变的函数 $F$,将 $n$ 个此类P-变量映射为一个单一的有效汇总P-变量 $V$。
  • 确定最小的递增函数 $f_{n,k}$,使得 $f_{n,k}(U_{k:n})$ 是一个有效的P-变量,其中 $U_{k:n}$ 是 $U_i$ 的第k阶统计量。
  • 为在弱依赖场景(如时间序列的子采样数据或分组观测)中组合P-值,提供一种实用且理论基础坚实的组合方法。

提出的方法

  • 该方法聚焦于第k阶统计量 $U_{k:n}$ 的函数,假设 $U_i$ 在给定数据条件下个体有效且条件独立同分布。
  • 定义 $f_{n,k}(u)$ 为满足在给定假设下 $f_{n,k}(U_{k:n})$ 是有效P-变量的最小递增函数。
  • 通过二项分布和顺序统计量的性质推导出解,特别利用了二项分布 $\mathrm{B}_{n,p}$ 的累积分布函数。
  • 通过涉及乘积测度和像测度的测度论论证,建立了关键不等式 $f_{n,k}(u) \leq 1 \land \left(\frac{n}{k}u\right)$。
  • 使用P-核和像测度的概念,形式化了在原假设下汇总P-变量有效性的理论基础。
  • 通过构造性证明表明,$V = 1 \land \left(\frac{n}{k}U_{k:n}\right)$ 始终是有效的P-变量,且在 $U_{k:n}$ 的保序函数类中具有最优性。

实验结果

研究问题

  • RQ1在 $U_1,\ldots,U_n$ 个体有效且条件独立同分布的假设下,使得 $f_{n,k}(U_{k:n})$ 为有效P-变量的最小递增函数 $f_{n,k}$ 是什么?
  • RQ2能否构造一个在弱依赖假设(如组内依赖)下既有效又稳健的汇总P-变量?
  • RQ3k 的选择如何影响组合P-变量的统计功效和保守性?
  • RQ4是否存在一种系统性方法来选择k,而非仅依赖于中位数等启发式选择?
  • RQ5在特定统计模型中,是否存在优于基于单一顺序统计量的汇总P-变量?

主要发现

  • 在所述假设下,函数 $f_{n,k}(u) \leq 1 \land \left(\frac{n}{k}u\right)$ 是使得 $f_{n,k}(U_{k:n})$ 为有效P-变量的最小递增函数。
  • 汇总P-变量 $V = 1 \land \left(\frac{n}{k}U_{k:n}\right)$ 始终有效,且当 $k$ 较大时该界是紧的。
  • 当 $k = \left\lfloor \frac{n+1}{2} \right\rfloor$ 时,$V$ 近似为 $U_i$ 左侧样本中位数的两倍,此选择被建议作为实际应用的默认值。
  • 通过涉及像测度和二项尾部概率的测度论论证,证明了 $f_{n,k}$ 的最优性。
  • 该结果在最弱假设下成立:即 $U_i$ 的个体有效性及其在给定数据下的条件独立同分布结构。
  • 该方法具有稳健性和分布无关性,对潜在数据生成过程无需任何参数假设。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。