[Paper Review] A gravity model for inter-city telephone communication networks
This study proposes a gravity model to explain inter-city telephone communication patterns using anonymized call data from 2.5 million Belgian mobile users. By aggregating users by zip code, the authors show that communication intensity between cities scales with the product of their populations divided by the square of their distance, confirming a gravity law similar to physical forces.
We consider a network of mobile phone customers aggregated by geographical proximity. We analyze the anonymous communication patterns of 2.5 million customers of a Belgian mobile phone operator. Grouping customers by billing address, we build a social network of cities, that consists of communication between 571 cities in Belgium. We show that inter-city communication intensity is characterized by a gravity model: the communication intensity between two cities is proportional to the product of the size of the population of these cities divided by the square of their distance. PACS numbers: 89.75.Da, 89.75.Fb, 89.65.Ef A gravity model for inter-city telephone communication networks 2 Recent research has shown that certain characteristics of cities grow in different ways in relation to population size. While some characteristics are directly proportional to a cities population size, instead, other features such as productivity or energy consumption are not linear but exhibit superlinear or sublinear dependence to population size [1]. Interestingly, some of these features have strong similarities with those found in biological cells an observation that has led to the creation of a metaphor where cities are seen as living entities [2]. Interactions between cities, such as passenger transport flows and phone messages, have also been related to population and distance [3, 4]. Meanwhile, in socio-economic networks, interactions between entities such as cities or countries have led to models remembering Newton’s gravity law, where the sizes of the entities play the role of mass [5]. Road and airline networks between cities have also been studied [6, 7], and in the case of road networks, it appears that the strength of interaction also follows a gravity law. While these results have provided a better understanding of the way cities interact, a finer analysis at human level was until now difficult because of a lack of data. Recently, however, telephone communication data has opened up a new way of analyzing cities at both a fine and aggregate level, whereby as Gottman as already as 1957 noted [8]: “the density of the flow of telephone calls is a fairly good measure of the relationships binding together the economic interests of the region”. Several large datasets of email and phone calls have recently become available. By using these as a proxy for social networks, they have enabled the study of human connections and behaviors [9, 10, 11, 12, 13]. The use of geographical information makes it possible to go one step further in the study of individual and group interactions. For example, Lambiotte et al. use a mobile phone dataset to show that the probability for a call between two people decreases by the square of their distance [14]. However, while the structure of complex networks has already been widely studied [15, 16, 17, 18], to date, contributions have not yet analyzed large-scale features of social networks where people are aggregated based on their geographical proximity. In this work, we study anonymized mobile phone communications from a Belgian operator and derive a model of interaction between cities. Grouping customers together by billing address, we create a two-level network, containing both a microscopic network of human-to-human interactions, and a macroscopic network of interactions between cities. The data that we consider consists of the communications made by more than 2.5 million customers of a Belgian mobile phone operator over a period of 6 months in 2006 [14]. Every customer is identified by a surrogate key and to every customer we associate their corresponding billing address zip code. In order to construct the communication network, we have filtered out calls involving other operators (there are three main operators in Belgium), incoming or outgoing, and we have kept only those transactions in which both the calling and receiving individuals are customers of the mobile phone company. In order to eliminate “accidental calls”, we have kept links between two customers i and j only if there are at least six calls in both directions A gravity model for inter-city telephone communication networks 3 during the 6 months time interval. The resulting network is composed of 2.5 million nodes and 38 million links. To the link between the customers i and j we associate a communication intensity by computing the total communication time in seconds lij between i and j. In order to analyze the relationship between this social network and geographical positioning, we associate customers to cities based on their billing address zip code. Belgium is a country of approximately 10.5 million inhabitants, with a high population density of 344 inhab./km. The Belgian National Institute of Statistics (NIS) [19] divides this population into 571 cities (cities, towns and villages), whose sizes show an overall lognormal population distribution with approximate parameters μ = 4.05 and σ = 0.37. ‡ The analyzed communication network provides information for the operator’s Figure 1. Ranks of city population sizes (blue triangles) and number of customers (red squares) follow similar distributions. customers rather than for the entire population. However, the number of customers present in each city follows the same lognormal distribution as the total population and so this suggests that our dataset is not structurally biased by particular user-groups and market shares. This is also confirmed by the ranks of city population sizes that match with those of customers, as shown in Fig. 1. In the rest of this article, when we use the term population of a city, we are refering to the number of customers that have a valid ZIP code of this city, even if they do not make calls during the six months period. There are a significant fraction of nodes that do not make any calls over the whole period, that are isolated nodes in the graph. These nodes are still taken into account for the population size, since their presence is of interest for the normalization of the communication data. By aggregating the individual communications at a city level, we obtain a network of 571 cities in Belgium. We define the intensity of interaction between the cities A and B by (Fig. 2 (a)): LAB = ∑
Motivation & Objective
- To model inter-city communication patterns using large-scale anonymized mobile phone data.
- To investigate whether communication intensity between cities follows a gravity law similar to physical forces.
- To analyze the relationship between city population size, geographic distance, and communication volume at scale.
- To validate the model using real-world mobile network data from Belgium over a 6-month period.
Proposed method
- Aggregated anonymized call records from 2.5 million Belgian mobile users, filtered to include only calls between customers of the same operator.
- Grouped users by zip code to define cities, with city population defined as the number of customers per city.
- Constructed a city-level communication network by summing total call duration between all city pairs.
- Applied a gravity model where communication intensity LAB is proportional to (PA × PB) / d²AB, with PA and PB as city populations and dAB as distance.
- Filtered links to include only bidirectional communication with at least six calls in each direction over six months to reduce noise.
- Used lognormal distribution to characterize city population sizes and validated data representativeness via customer rank correlation.
Experimental results
Research questions
- RQ1Does inter-city communication intensity in a national mobile network follow a gravity law based on population size and distance?
- RQ2How well does the product of city populations divided by squared distance predict actual communication volume between cities?
- RQ3To what extent is the observed communication pattern independent of user-specific behavior and representative of the broader population?
- RQ4Are the communication patterns robust to filtering for sustained interaction (e.g., six calls in each direction)?
- RQ5Can large-scale anonymized mobile data reliably model macroscopic urban interaction patterns?
Key findings
- Inter-city communication intensity between cities in Belgium follows a gravity model, with intensity proportional to the product of city populations divided by the square of their distance.
- The model explains a significant portion of the variance in communication volume across city pairs, indicating strong predictive power.
- City population size, as measured by customer count per zip code, follows a lognormal distribution with parameters μ = 4.05 and σ = 0.37.
- The rank ordering of city populations closely matches the rank ordering of customer counts, suggesting no structural bias in the dataset.
- The filtered network includes 38 million links among 2.5 million users, with 571 cities forming the macroscopic network.
- Isolated nodes (users with no calls) are retained in population counts, ensuring accurate normalization of communication intensity.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.