[论文解读] A Deep Generative Model for Feasible and Diverse Population Synthesis
该论文提出一种带有定制正则化的深度生成模型,用于人口合成,旨在最小化结构零(不可行的属性组合),同时保留抽样零(真实人口中存在但未在调查样本中观测到的有效组合)。通过将这些正则化应用于变分自编码器(VAE)和生成对抗网络(GAN),该方法在可行性与多样性之间取得平衡,分别实现79.2%的精确率(20.8%的结构零)和89.0%的精确率(11.0%的结构零),优于传统模型。
An ideal synthetic population, a key input to activity-based models, mimics the distribution of the individual- and household-level attributes in the actual population. Since the entire population's attributes are generally unavailable, household travel survey (HTS) samples are used for population synthesis. Synthesizing population by directly sampling from HTS ignores the attribute combinations that are unobserved in the HTS samples but exist in the population, called 'sampling zeros'. A deep generative model (DGM) can potentially synthesize the sampling zeros but at the expense of generating 'structural zeros' (i.e., the infeasible attribute combinations that do not exist in the population). This study proposes a novel method to minimize structural zeros while preserving sampling zeros. Two regularizations are devised to customize the training of the DGM and applied to a generative adversarial network (GAN) and a variational autoencoder (VAE). The adopted metrics for feasibility and diversity of the synthetic population indicate the capability of generating sampling and structural zeros -- lower structural zeros and lower sampling zeros indicate the higher feasibility and the lower diversity, respectively. Results show that the proposed regularizations achieve considerable performance improvement in feasibility and diversity of the synthesized population over traditional models. The proposed VAE additionally generated 23.5% of the population ignored by the sample with 79.2% precision (i.e., 20.8% structural zeros rates), while the proposed GAN generated 18.3% of the ignored population with 89.0% precision. The proposed improvement in DGM generates a more feasible and diverse synthetic population, which is critical for the accuracy of an activity-based model.
研究动机与目标
- 为解决在依赖有限家庭出行调查(HTS)样本时,生成既可行又多样的合成人口的挑战。
- 最小化深度生成模型在人口合成过程中产生的结构零——即无效或不可行的属性组合。
- 保留抽样零——即在HTS样本中未观测到但在真实人口中有效的属性组合,从而提升合成人口的代表性。
- 通过生成更真实、更全面的合成人口,提升基于活动的模型的准确性。
提出的方法
- 提出两种新颖的正则化方法,以指导深度生成模型(VAE和GAN)的训练,减少结构零的同时保留抽样零。
- 在GAN的对抗训练过程中以及VAE的潜在空间优化过程中应用正则化,以强制实施可行性约束。
- 使用一种具备可行性意识的损失函数,基于已知的人口约束,对生成不可行属性组合的行为施加惩罚。
- 引入一种具备多样性意识的度量方法,用于评估未在HTS样本中观测到但有效的稀有属性组合的保留程度。
- 采用两阶段训练流程:首先在HTS数据上预训练模型;其次应用正则化以纠正不可行和缺失的组合。
- 通过可行性(结构零率)和多样性(抽样零保留率)的度量验证模型性能,并在真实世界的HTS数据上进行定量评估。
实验结果
研究问题
- RQ1能否对深度生成模型进行正则化,以在人口合成中减少结构零,同时保留抽样零?
- RQ2与传统方法相比,所提出的正则化在提升合成人口的可行性和多样性方面效果如何?
- RQ3VAE或GAN在多大程度上能够生成之前未被观测到但在人口中有效的属性组合?
- RQ4在合成人口生成中,可行性和多样性之间的权衡如何?应如何优化?
- RQ5在真实世界的HTS数据中,所提出的正则化对合成人口生成的精确率和召回率有何影响?
主要发现
- 所提出的VAE生成了HTS样本忽略的23.5%的人口,精确率为79.2%,对应20.8%的结构零率。
- 所提出的GAN生成了被HTS样本忽略的18.3%的人口,精确率为89.0%,表明结构零率为11.0%。
- 两种模型在平衡可行性和多样性方面显著优于传统的人口合成方法。
- 正则化方法成功减少了结构零,同时保持了抽样零的高水平保留,提升了合成人口的真实性。
- VAE在召回未被观测但有效的组合方面表现更优,而GAN在精确率方面表现更佳,体现了互补优势。
- 结果证实,所提出的方法能够生成更准确、更全面的合成人口,这对基于活动的建模至关重要。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。