Skip to main content
QUICK REVIEW

[论文解读] Semigroups and sequential importance sampling for multiway tables

Ruriko Yoshida, Jing Xi|arXiv (Cornell University)|Nov 28, 2011
Advanced Combinatorial Mathematics参考文献 13被引用 3
一句话总结

本文提出了一种多项式时间的顺序重要性抽样(SIS)算法,用于在不拒绝的情况下、在固定边际约束下、基于均匀分布生成列联表。证明了当设计矩阵规模固定时,SIS 可以高效地抽样有效表格,尽管在实际中拒绝率仍然很高,尤其是在逻辑斯蒂回归模型下。

ABSTRACT

When an interval of integers between the lower bound $l_i$ and the upper bound $u_i$ is the support of the marginal distribution $n_i|(n_{i-1}, ...,n_1)$, Chen et al, 2005 noticed that sampling from the interval at each step, for $n_i$ during a sequential importance sampling (SIS) procedure, always produces a table which satisfies the marginal constraints. However, in general, the interval may not be equal to the support of the marginal distribution. In this case, the SIS procedure may produce tables which do not satisfy the marginal constraints, leading to rejection Chen et al 2006. In this paper we consider the uniform distribution as the target distribution. First we show that if we fix the number of rows and columns of the design matrix of the model for contingency tables then there exists a polynomial time algorithm in terms of the input size to sample a table from the set of all tables satisfying all marginals defined by the given model via the SIS procedure without rejection. We then show experimentally that in general the SIS procedure may have large rejection rates even with small tables. Further we show that in general the classical SIS procedure in Chen et al, 2005 can have a large rejection rate whose limit is one. When estimating the number of tables in our simulation study, we used the univariate and bivariate logistic regression models since under this model the SIS procedure seems to have higher rate of rejections even with small tables.

研究动机与目标

  • 开发一种在固定边际约束下抽样列联表的无拒绝SIS程序。
  • 建立SIS在无拒绝情况下高效运行的条件,特别是当设计矩阵规模固定时。
  • 研究经典SIS方法在实际中的拒绝率,特别是在逻辑斯蒂回归模型下的情况。
  • 评估SIS在不同模型假设下估计列联表数量时的性能。

提出的方法

  • 使用顺序重要性抽样(SIS)逐步生成多维列联表,基于先前条目的条件分布。
  • 将均匀分布作为目标分布,用于抽样满足所有边际约束的表格。
  • 证明当设计矩阵的行数和列数固定时,SIS 可以在多项式时间内无拒绝地运行。
  • 采用单变量和双变量逻辑斯蒂回归模型,模拟并分析SIS中的拒绝率。
  • 分析条件分布的支持集,以确定在边界内抽样是否能确保边际约束的满足。
  • 通过模拟研究评估在不同模型设定下拒绝率和抽样效率。

实验结果

研究问题

  • RQ1在何种条件下,顺序重要性抽样能够无拒绝地生成满足所有边际约束的列联表?
  • RQ2当设计矩阵规模固定时,SIS 抽样具有固定边际的列联表的计算复杂度是多少?
  • RQ3即使对于小型列联表,经典SIS程序的拒绝率有多高?
  • RQ4模型选择(如逻辑斯蒂回归)是否显著影响SIS在列联表抽样中的拒绝率?
  • RQ5能否为具有固定边际的列联表的均匀抽样构建一种无拒绝的SIS算法?

主要发现

  • 当设计矩阵的行数和列数固定时,存在一种多项式时间SIS算法,可无拒绝地抽样有效列联表。
  • 经典SIS程序(Chen et al., 2005)即使在小型表格中,其拒绝率在极限下也可能趋近于1。
  • 即使在小型表格中,使用逻辑斯蒂回归模型作为基础模型时,也观察到较高的拒绝率。
  • 条件分布的支持集并不总等于由边际边界定义的区间,这可能导致标准SIS中违反约束。
  • 模拟结果证实,与简单模型相比,逻辑斯蒂回归模型下的拒绝率显著更高。
  • 本文建立了SIS可实现无拒绝的理论条件,为高效抽样算法提供了基础。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。