[Paper Review] NWPU-Crowd: A Large-Scale Benchmark for Crowd Counting.
This paper introduces NWPU-Crowd, a large-scale benchmark with 5,109 images and 2.13 million annotated heads for crowd counting and localization. It features extreme density variation (0–20,033) and diverse lighting conditions, enabling rigorous evaluation of CNN-based methods through a public benchmark website, significantly advancing the state of the art in crowd counting research.
In the last decade, crowd counting and localization attract much attention of researchers due to its wide-spread applications, including crowd monitoring, public safety, space design, etc. Many Convolutional Neural Networks (CNN) are designed for tackling this task. However, currently released datasets are so small-scale that they can not meet the needs of the supervised CNN-based algorithms. To remedy this problem, we construct a large-scale congested crowd counting and localization dataset, NWPU-Crowd, consisting of 5,109 images, in a total of 2,133,375 annotated heads with points and boxes. Compared with other real-world datasets, it contains various illumination scenes and has the largest density range (0~20,033). Besides, a benchmark website is developed for impartially evaluating the different methods, which allows researchers to submit the results of the test set. Based on the proposed dataset, we further describe the data characteristics, evaluate the performance of some mainstream state-of-the-art (SOTA) methods, and analyze the new problems that arise on the new data. What's more, the benchmark is deployed at \url{this https URL}, and the dataset/code/models/results are available at \url{this https URL}.
Motivation & Objective
- To address the limitation of existing crowd counting datasets, which are too small-scale to support training and evaluating deep CNN-based methods.
- To provide a large-scale, diverse, and realistic dataset with extreme density variation and varied illumination conditions.
- To establish a public benchmark website for impartial evaluation of crowd counting methods on a standardized test set.
- To analyze new challenges arising in high-density, complex scenes using the proposed dataset.
- To facilitate the development and comparison of state-of-the-art crowd counting models through shared data, code, and results.
Proposed method
- Construction of a large-scale crowd counting dataset, NWPU-Crowd, comprising 5,109 real-world images with point and box annotations for individual heads.
- Annotation of 2,133,375 heads across diverse scenes, including extreme congestion and varying lighting conditions.
- Design of a benchmark website to host the test set and enable standardized, impartial evaluation of submitted methods.
- Collection of data with the largest reported density range (0 to 20,033) to challenge existing models.
- Evaluation of multiple state-of-the-art (SOTA) crowd counting models on the new dataset to identify performance gaps and new failure modes.
- Deployment of the dataset, code, models, and results at a public URL for open access and reproducibility.
Experimental results
Research questions
- RQ1How do existing state-of-the-art crowd counting models perform on a large-scale, high-density, and diverse real-world dataset like NWPU-Crowd?
- RQ2What new challenges emerge in extreme density and complex illumination conditions that are not adequately addressed by current models?
- RQ3How does the proposed benchmark enable fair and reproducible evaluation of crowd counting methods?
- RQ4What are the key data characteristics of NWPU-Crowd that differentiate it from existing datasets in terms of scale, diversity, and density range?
- RQ5How does the inclusion of both point and box annotations enhance the evaluation and localization capability of crowd counting models?
Key findings
- The NWPU-Crowd dataset contains 5,109 images with 2,133,375 annotated heads, representing the largest and most diverse crowd counting dataset to date.
- The dataset spans an unprecedented density range of 0 to 20,033, exposing limitations in current models trained on smaller, less dense data.
- Evaluation on the benchmark reveals significant performance drops for several SOTA models in high-density and low-illumination scenarios, indicating new failure modes.
- The benchmark website enables standardized and impartial evaluation, promoting reproducibility and fair comparison across methods.
- The dataset and benchmark facilitate the discovery of new challenges in crowd counting, such as severe occlusion and extreme density variation.
- The availability of code, models, and results at a public URL enhances transparency and accelerates research progress in the field.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.