[Paper Review] Developing synthetic individual-level population datasets: The case of contextualizing maps of privacy-preserving census data
This paper presents a method to generate open, realistic synthetic individual-level population datasets by optimizing population distributions to match published census statistics, enabling contextualization of maps created with privacy-preserving census data. Using two Ohio counties as case studies, the approach minimizes discrepancies between synthetic and real census data, demonstrating its utility in teaching and analyzing cartographic artifacts in privacy-protected data environments.
The purpose of this paper is to describe the development of a synthetic population dataset that is open and realistic and can be used to facilitate understanding the cartographic process and contextualizing the cartographic artifacts. We first discuss an optimization model that is designed to construct the synthetic population by minimizing the difference between the summarized information of the synthetic populations and the statistics published in census data tables. We then illustrate how the synthetic population dataset can be used to contextualize maps made using privacy-preserving census data. Two counties in Ohio are used as case studies.
Motivation & Objective
- To develop a realistic, open-access synthetic individual-level population dataset that mirrors real census statistics.
- To enable better understanding of cartographic processes and artifacts in privacy-preserving data environments.
- To reduce the gap between synthetic population data and published census statistics through optimization.
- To demonstrate the utility of synthetic data in contextualizing maps derived from privacy-protected census data.
- To provide a reproducible framework for generating synthetic populations that support geospatial education and research.
Proposed method
- An optimization model is formulated to minimize the difference between summarized statistics of the synthetic population and published census data tables.
- The model uses constraints derived from known census statistics to ensure synthetic individuals reflect real demographic distributions.
- The optimization process assigns individual-level attributes (e.g., age, race, housing) while preserving aggregate statistics.
- The synthetic dataset is generated at the individual level, enabling detailed spatial and demographic analysis.
- Case studies in two Ohio counties validate the model’s ability to reproduce real census patterns at the block group level.
- The method supports the creation of realistic maps using privacy-preserving data by providing a reference synthetic dataset for comparison.
Experimental results
Research questions
- RQ1How can a synthetic individual-level population dataset be generated to closely match published census statistics?
- RQ2To what extent can synthetic data replicate real demographic distributions in specific geographic areas?
- RQ3How can synthetic datasets improve the contextual understanding of maps created from privacy-preserving census data?
- RQ4What optimization techniques are effective in aligning synthetic population data with real census aggregates?
- RQ5Can synthetic data serve as a reliable benchmark for evaluating cartographic outputs from privacy-protected data?
Key findings
- The optimization model successfully generated a synthetic population dataset that closely aligns with published census statistics in two Ohio counties.
- The synthetic dataset enables accurate contextualization of maps produced using privacy-preserving census data, improving interpretability.
- The method effectively minimizes differences between synthetic and real census data at the block group level, demonstrating high fidelity.
- The synthetic dataset supports educational and analytical use by providing a realistic reference for cartographic artifacts.
- The approach is reproducible and scalable, offering a framework applicable to other geographic regions and data types.
- The results show that synthetic data can serve as a valuable proxy for understanding the implications of data suppression and perturbation in privacy-preserving census data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.