Skip to main content
QUICK REVIEW

[論文レビュー] A gravity model for inter-city telephone communication networks

Gautier Krings, Francesco Calabrese|arXiv (Cornell University)|May 5, 2009
Complex Network Analysis Techniques参考文献 22被引用数 15
ひとこと要約

本研究では、250万件のベルギーのモバイルユーザーの匿名化された通話データを用いて、都市間の電話通話パターンを説明する重力モデルを提案する。郵便番号ごとにユーザーを集約することで、都市間の通信強度が、人口の積を距離の二乗で割ったものに比例することを示し、物理的力に類似した重力法則が成り立つことを確認した。

ABSTRACT

We consider a network of mobile phone customers aggregated by geographical proximity. We analyze the anonymous communication patterns of 2.5 million customers of a Belgian mobile phone operator. Grouping customers by billing address, we build a social network of cities, that consists of communication between 571 cities in Belgium. We show that inter-city communication intensity is characterized by a gravity model: the communication intensity between two cities is proportional to the product of the size of the population of these cities divided by the square of their distance. PACS numbers: 89.75.Da, 89.75.Fb, 89.65.Ef A gravity model for inter-city telephone communication networks 2 Recent research has shown that certain characteristics of cities grow in different ways in relation to population size. While some characteristics are directly proportional to a cities population size, instead, other features such as productivity or energy consumption are not linear but exhibit superlinear or sublinear dependence to population size [1]. Interestingly, some of these features have strong similarities with those found in biological cells an observation that has led to the creation of a metaphor where cities are seen as living entities [2]. Interactions between cities, such as passenger transport flows and phone messages, have also been related to population and distance [3, 4]. Meanwhile, in socio-economic networks, interactions between entities such as cities or countries have led to models remembering Newton’s gravity law, where the sizes of the entities play the role of mass [5]. Road and airline networks between cities have also been studied [6, 7], and in the case of road networks, it appears that the strength of interaction also follows a gravity law. While these results have provided a better understanding of the way cities interact, a finer analysis at human level was until now difficult because of a lack of data. Recently, however, telephone communication data has opened up a new way of analyzing cities at both a fine and aggregate level, whereby as Gottman as already as 1957 noted [8]: “the density of the flow of telephone calls is a fairly good measure of the relationships binding together the economic interests of the region”. Several large datasets of email and phone calls have recently become available. By using these as a proxy for social networks, they have enabled the study of human connections and behaviors [9, 10, 11, 12, 13]. The use of geographical information makes it possible to go one step further in the study of individual and group interactions. For example, Lambiotte et al. use a mobile phone dataset to show that the probability for a call between two people decreases by the square of their distance [14]. However, while the structure of complex networks has already been widely studied [15, 16, 17, 18], to date, contributions have not yet analyzed large-scale features of social networks where people are aggregated based on their geographical proximity. In this work, we study anonymized mobile phone communications from a Belgian operator and derive a model of interaction between cities. Grouping customers together by billing address, we create a two-level network, containing both a microscopic network of human-to-human interactions, and a macroscopic network of interactions between cities. The data that we consider consists of the communications made by more than 2.5 million customers of a Belgian mobile phone operator over a period of 6 months in 2006 [14]. Every customer is identified by a surrogate key and to every customer we associate their corresponding billing address zip code. In order to construct the communication network, we have filtered out calls involving other operators (there are three main operators in Belgium), incoming or outgoing, and we have kept only those transactions in which both the calling and receiving individuals are customers of the mobile phone company. In order to eliminate “accidental calls”, we have kept links between two customers i and j only if there are at least six calls in both directions A gravity model for inter-city telephone communication networks 3 during the 6 months time interval. The resulting network is composed of 2.5 million nodes and 38 million links. To the link between the customers i and j we associate a communication intensity by computing the total communication time in seconds lij between i and j. In order to analyze the relationship between this social network and geographical positioning, we associate customers to cities based on their billing address zip code. Belgium is a country of approximately 10.5 million inhabitants, with a high population density of 344 inhab./km. The Belgian National Institute of Statistics (NIS) [19] divides this population into 571 cities (cities, towns and villages), whose sizes show an overall lognormal population distribution with approximate parameters μ = 4.05 and σ = 0.37. ‡ The analyzed communication network provides information for the operator’s Figure 1. Ranks of city population sizes (blue triangles) and number of customers (red squares) follow similar distributions. customers rather than for the entire population. However, the number of customers present in each city follows the same lognormal distribution as the total population and so this suggests that our dataset is not structurally biased by particular user-groups and market shares. This is also confirmed by the ranks of city population sizes that match with those of customers, as shown in Fig. 1. In the rest of this article, when we use the term population of a city, we are refering to the number of customers that have a valid ZIP code of this city, even if they do not make calls during the six months period. There are a significant fraction of nodes that do not make any calls over the whole period, that are isolated nodes in the graph. These nodes are still taken into account for the population size, since their presence is of interest for the normalization of the communication data. By aggregating the individual communications at a city level, we obtain a network of 571 cities in Belgium. We define the intensity of interaction between the cities A and B by (Fig. 2 (a)): LAB = ∑

研究の動機と目的

  • 大規模な匿名化されたモバイル電話データを用いて都市間の通信パターンをモデル化すること。
  • 都市間の通信強度が、物理的力に類似した重力法則に従うかどうかを調査すること。
  • 都市の人口規模、地理的距離、および通信量の関係を大規模スケールで分析すること。
  • ベルギーの6か月間の実際のモバイルネットワークデータを用いて、モデルを検証すること。

提案手法

  • 同一の通信事業者を利用している顧客間の通話に限定した、250万件のベルギーのモバイルユーザーの集約化された匿名通話記録。
  • 都市を郵便番号ごとにグループ化し、都市の人口を各都市の顧客数として定義。
  • すべての都市ペア間の合計通話時間の合計をもとに、都市レベルの通信ネットワークを構築。
  • 通信強度 LAB が (PA × PB) / d²AB に比例する重力モデルを適用、ここで PA と PB は都市の人口、dAB は距離である。
  • ノイズを低減するために、6か月間にわたり片方向に少なくとも6通話以上の双方向通信を含むリンクにフィルタリング。
  • 都市の人口サイズを対数正規分布で特徴付け、顧客のランク相関を用いてデータの代表性を検証。

実験結果

リサーチクエスチョン

  • RQ1国家規模のモバイルネットワークにおける都市間の通信強度は、人口規模と距離に基づいた重力法則に従うか?
  • RQ2都市の人口の積を距離の二乗で割った値が、実際の都市間通信量をどれほどよく予測できるか?
  • RQ3観察された通信パターンは、ユーザー固有の行動に依存せず、広範な人口の代表的パターンであると見なせるか?
  • RQ4継続的な相互作用(例:片方向に6通話以上)をフィルタリングした場合でも、通信パターンは頑健であるか?
  • RQ5大規模な匿名化されたモバイルデータは、顕著な都市相互作用パターンを信頼性高くモデル化できるか?

主な発見

  • ベルギーの都市間通信強度は、人口の積を距離の二乗で割ったものに比例する重力モデルに従うことが確認された。
  • モデルは都市ペア間の通信量の分散の大部分を説明でき、強い予測力を持つことが示された。
  • 郵便番号ごとの顧客数として測定された都市の人口規模は、パラメータ μ = 4.05 と σ = 0.37 の対数正規分布に従う。
  • 都市の人口ランク順序と顧客数のランク順序が非常に近いことから、データセットに構造的バイアスがないことが示された。
  • フィルタリングされたネットワークには、250万ユーザーの間で3800万のリンクがあり、571の都市がマクロスコピックなネットワークを形成した。
  • 通話のない孤立ノード(ユーザー)は、人口カウントに保持され、通信強度の正規化が正確に行われた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。