Skip to main content
QUICK REVIEW

[論文レビュー] An empirical study on the names of points of interest and their changes with geographic distance

Yingjie Hu, Krzysztof Janowicz|arXiv (Cornell University)|Jun 21, 2018
Geographic Information Systems Studies参考文献 17被引用数 6
ひとこと要約

本研究では、米国7都市圏のヤフー・データを用いて112,071件のポイント・オブ・インタレスト(POI)を分析し、POI名の用語が地理的距離とともにどのように変化するかを検討した。TF-IDFおよびword2vecを用いて現地語を同定し、意味的類似度を測定した結果、名前の類似度に顕著な距離減衰効果が認められ、勾配は-0.090、決定係数は0.828であった。これは、地理的距離が増加するにつれてPOI名の類似度が低下することを示している。

ABSTRACT

While Points Of Interest (POIs), such as restaurants, hotels, and barber shops, are part of urban areas irrespective of their specific locations, the names of these POIs often reveal valuable information related to local culture, landmarks, influential families, figures, events, and so on. Place names have long been studied by geographers, e.g., to understand their origins and relations to family names. However, there is a lack of large-scale empirical studies that examine the localness of place names and their changes with geographic distance. In addition to enhancing our understanding of the coherence of geographic regions, such empirical studies are also significant for geographic information retrieval where they can inform computational models and improve the accuracy of place name disambiguation. In this work, we conduct an empirical study based on 112,071 POIs in seven US metropolitan areas extracted from an open Yelp dataset. We propose to adopt term frequency and inverse document frequency in geographic contexts to identify local terms used in POI names and to analyze their usages across different POI types. Our results show an uneven usage of local terms across POI types, which is highly consistent among different geographic regions. We also examine the decaying effect of POI name similarity with the increase of distance among POIs. While our analysis focuses on urban POI names, the presented methods can be generalized to other place types as well, such as mountain peaks and streets.

研究の動機と目的

  • POI名の用語が地域文化をどのように反映し、地理空間においてどのように変化するかを理解すること。
  • 地理的TF-IDFを用いてPOI名の現地語を同定し、POIの種別ごとの使用状況を評価すること。
  • 地域間のPOI名の集団的意味的類似度をモデル化し、距離減衰効果の有無を検証すること。
  • 地理的情報検索および場所名の意味の解釈のための定量的基盤を提供すること。

提案手法

  • 米国7都市圏にまたがるオープンなヤフー・データセットから112,071件のPOIを抽出した。
  • 地理的文脈における項目の頻度・逆文書頻度(TF-IDF)を適用し、POI名に含まれる現地語を同定した。
  • word2vec埋め込みを用いて、地域間のPOI名の意味的類似度をモデル化した。
  • カウントベースのベクトルとword2vec埋め込みの両方を用いて、都市圏間の集団的類似度行列を計算した。
  • 地理的距離の増加に伴う名前類似度の減衰を評価するため、線形モデルをフィットさせた。
  • 類似度パターンと地理的距離行列の可視化を用いて、結果の妥当性を検証した。

実験結果

リサーチクエスチョン

  • RQ1POIの種別ごとに、POI名の現地語はどのように変化するのか。このパターンは地理的地域間で一貫しているか。
  • RQ2都市圏間の地理的距離が増加するにつれて、POI名の類似度はどの程度低下するか。
  • RQ3カウントベースのベクトル表現とword2vec埋め込みは、POI名の意味的類似度をどの程度正確に捉えているか。
  • RQ4地理的距離と地域間のPOI名の集団的類似度の間には、定量的な関係が存在するか。
  • RQ5観察された現地語の使用パターンと名前の類似度パターンは、場所名の意味の解釈のための計算モデルにどのように活用できるか。

主な発見

  • POI名の現地語は、頻度分布においてジプフの法則に従っており、用語使用にパワー・ルールのパターンが見られる。
  • 現地語の使用はPOIの種別によって不均一であり、自動車整備サービスではレストランよりも高い現地語使用率を示した。
  • POI名の類似度には顕著な距離減衰効果が認められ、word2vec埋め込みを用いた場合、勾配は-0.090、決定係数は0.828であった。
  • word2vecベースの類似度測定値は、カウントベースのベクトルよりも地理的距離のパターンとよりよく一致した。
  • 類似度パターンから、地理的に近いとされる都市圏ペアと、遠いとされるペアに対応する2つの明確なクラスタが特定された。
  • 結果から、POI名の意味論が地理的接近性によって体系的に影響を受けることが示され、地理情報システムにおける現地語特徴の活用が有効であることが支持された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。