[논문 리뷰] A gravity model for inter-city telephone communication networks
이 연구는 벨기에 모바일 사용자 250만 명의 익명화된 통화 데이터를 바탕으로 도시 간 통화 통신 패턴을 설명하기 위해 중력 모델을 제안한다. 사용자를 우편번호로 집계함으로써, 도시 간 통신 강도가 그들의 인구 수의 곱을 거리의 제곱으로 나눈 값에 비례함을 보여주며, 물리적 힘과 유사한 중력 법칙을 확인한다.
We consider a network of mobile phone customers aggregated by geographical proximity. We analyze the anonymous communication patterns of 2.5 million customers of a Belgian mobile phone operator. Grouping customers by billing address, we build a social network of cities, that consists of communication between 571 cities in Belgium. We show that inter-city communication intensity is characterized by a gravity model: the communication intensity between two cities is proportional to the product of the size of the population of these cities divided by the square of their distance. PACS numbers: 89.75.Da, 89.75.Fb, 89.65.Ef A gravity model for inter-city telephone communication networks 2 Recent research has shown that certain characteristics of cities grow in different ways in relation to population size. While some characteristics are directly proportional to a cities population size, instead, other features such as productivity or energy consumption are not linear but exhibit superlinear or sublinear dependence to population size [1]. Interestingly, some of these features have strong similarities with those found in biological cells an observation that has led to the creation of a metaphor where cities are seen as living entities [2]. Interactions between cities, such as passenger transport flows and phone messages, have also been related to population and distance [3, 4]. Meanwhile, in socio-economic networks, interactions between entities such as cities or countries have led to models remembering Newton’s gravity law, where the sizes of the entities play the role of mass [5]. Road and airline networks between cities have also been studied [6, 7], and in the case of road networks, it appears that the strength of interaction also follows a gravity law. While these results have provided a better understanding of the way cities interact, a finer analysis at human level was until now difficult because of a lack of data. Recently, however, telephone communication data has opened up a new way of analyzing cities at both a fine and aggregate level, whereby as Gottman as already as 1957 noted [8]: “the density of the flow of telephone calls is a fairly good measure of the relationships binding together the economic interests of the region”. Several large datasets of email and phone calls have recently become available. By using these as a proxy for social networks, they have enabled the study of human connections and behaviors [9, 10, 11, 12, 13]. The use of geographical information makes it possible to go one step further in the study of individual and group interactions. For example, Lambiotte et al. use a mobile phone dataset to show that the probability for a call between two people decreases by the square of their distance [14]. However, while the structure of complex networks has already been widely studied [15, 16, 17, 18], to date, contributions have not yet analyzed large-scale features of social networks where people are aggregated based on their geographical proximity. In this work, we study anonymized mobile phone communications from a Belgian operator and derive a model of interaction between cities. Grouping customers together by billing address, we create a two-level network, containing both a microscopic network of human-to-human interactions, and a macroscopic network of interactions between cities. The data that we consider consists of the communications made by more than 2.5 million customers of a Belgian mobile phone operator over a period of 6 months in 2006 [14]. Every customer is identified by a surrogate key and to every customer we associate their corresponding billing address zip code. In order to construct the communication network, we have filtered out calls involving other operators (there are three main operators in Belgium), incoming or outgoing, and we have kept only those transactions in which both the calling and receiving individuals are customers of the mobile phone company. In order to eliminate “accidental calls”, we have kept links between two customers i and j only if there are at least six calls in both directions A gravity model for inter-city telephone communication networks 3 during the 6 months time interval. The resulting network is composed of 2.5 million nodes and 38 million links. To the link between the customers i and j we associate a communication intensity by computing the total communication time in seconds lij between i and j. In order to analyze the relationship between this social network and geographical positioning, we associate customers to cities based on their billing address zip code. Belgium is a country of approximately 10.5 million inhabitants, with a high population density of 344 inhab./km. The Belgian National Institute of Statistics (NIS) [19] divides this population into 571 cities (cities, towns and villages), whose sizes show an overall lognormal population distribution with approximate parameters μ = 4.05 and σ = 0.37. ‡ The analyzed communication network provides information for the operator’s Figure 1. Ranks of city population sizes (blue triangles) and number of customers (red squares) follow similar distributions. customers rather than for the entire population. However, the number of customers present in each city follows the same lognormal distribution as the total population and so this suggests that our dataset is not structurally biased by particular user-groups and market shares. This is also confirmed by the ranks of city population sizes that match with those of customers, as shown in Fig. 1. In the rest of this article, when we use the term population of a city, we are refering to the number of customers that have a valid ZIP code of this city, even if they do not make calls during the six months period. There are a significant fraction of nodes that do not make any calls over the whole period, that are isolated nodes in the graph. These nodes are still taken into account for the population size, since their presence is of interest for the normalization of the communication data. By aggregating the individual communications at a city level, we obtain a network of 571 cities in Belgium. We define the intensity of interaction between the cities A and B by (Fig. 2 (a)): LAB = ∑
연구 동기 및 목표
- 대규모 익명화된 모바일 휴대전화 데이터를 사용하여 도시 간 통신 패턴을 모델링하기 위해.
- 도시 간 통신 강도가 물리적 힘과 유사한 중력 법칙을 따르는지 조사하기 위해.
- 도시 인구 규모, 지리적 거리, 통신량 간의 관계를 대규모 스케일에서 분석하기 위해.
- 벨기에에서 6개월 간의 실제 모바일 네트워크 데이터를 사용하여 모델을 검증하기 위해.
제안 방법
- 동일한 통신사 고객 간 통화만 포함하도록 필터링된 250만 명의 벨기에 모바일 사용자로부터의 익명화된 통화 기록을 집계함.
- 사용자를 우편번호 기준으로 그룹화하여 도시를 정의하고, 도시의 인구를 각 도시의 고객 수로 정의함.
- 모든 도시 쌍 간의 총 통화 시간을 합산하여 도시 수준의 통신 네트워크를 구축함.
- 통신 강도 LAB가 (PA × PB) / d²AB 비례하는 중력 모델을 적용함. 여기서 PA와 PB는 도시의 인구 수이고, dAB는 거리임.
- 노이즈를 줄이기 위해 양방향 통화 중 각 방향으로 최소 6회 이상의 통화가 있었던 링크만 필터링함.
- 도시 인구 규모를 로그정규분포로 특성화하고, 고객 순위 상관관계를 통해 데이터의 대표성을 검증함.
실험 결과
연구 질문
- RQ1국가 차원의 모바일 네트워크에서 도시 간 통신 강도는 인구 규모와 거리에 기반한 중력 법칙을 따르는가?
- RQ2도시 인구 수의 곱을 거리의 제곱으로 나눈 값이 실제로 도시 간 통신량을 얼마나 잘 예측하는가?
- RQ3관찰된 통신 패턴이 사용자별 행동에 영향을 받지 않고 보편적 인구 집단을 대표하는가?
- RQ4지속적인 상호작용(예: 각 방향으로 최소 6회 통화)에 대해 필터링했을 때 통신 패턴이 얼마나 강인한가?
- RQ5대규모 익명화된 모바일 데이터로 거시적 도시 상호작용 패턴을 신뢰성 있게 모델링할 수 있는가?
주요 결과
- 벨기에의 도시 간 통신 강도는 중력 모델을 따르며, 강도가 도시 인구 수의 곱을 거리의 제곱으로 나눈 값에 비례함.
- 모델은 도시 쌍 간 통신량의 변동성의 상당 부분을 설명하며, 강력한 예측 능력을 보임.
- 우편번호 기반 고객 수로 측정된 도시 인구 규모는 μ = 4.05 및 σ = 0.37를 가진 로그정규분포를 따른다.
- 도시 인구 순위가 고객 수 순위와 밀도 있게 일치함을 통해 데이터셋에 구조적 편향이 없음을 시사함.
- 필터링된 네트워크에는 250만 명의 사용자 간 3,800만 개의 링크가 포함되어 있으며, 571개의 도시가 거시적 네트워크를 형성함.
- 통화가 전혀 없는 고립된 노드(사용자)는 인구 수 계산에 포함되며, 이는 통신 강도의 정규화를 정확히 하기 위함임.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.