[論文レビュー] Design and Implementation of iMacros-based Data Crawler for Behavioral Analysis of Facebook Users.
本稿では、法的境界内においてFacebookのプロフィールから個人情報およびウォール活動データを抽出するiMacrosベースのWebクローラーであるIMcrawlerを提示する。このシステムはデータを収集・前処理し、2つのデータセットに整理することで、情報開示における性別差異やオンライン活動のパターンを明らかにする。
Obtaining the desired dataset is still a prime challenge faced by researchers while analysing Online Social Network (OSN) sites. Application Programming Interfaces (APIs) provided by OSN service providers for retrieving data impose several unavoidable restrictions which make it difficult to get a desirable dataset. In this paper, we present an iMacros technology-based data crawler called IMcrawler,capable of collecting every piece of information which is accessible through a browser from the Facebook website within the legal framework reauthorized by Facebook.The proposed crawler addresses most of the challenges allied with web data extraction approaches and most of the APIs provided by OSN service providers. Two broad sections have been extracted from Facebook user profiles, namely, Personal Information and Wall Activities. The collected data is pre-processed into two datasets and each data set is statistically analysed to draw semantic knowledge and understand the several behavioral aspects of Facebook users such as kind of information mostly disclosed by users, gender differences in the pattern of revealed information, highly posted content on the network, highly performed activities on the network, the relationships among personal and post attributes, etc. To the best of our knowledge, the present work is the first attempt towards providing the detailed description of crawler design and gender-based information revealing behaviour of Facebook users.
研究の動機と目的
- オンラインソーシャルネットワーク(OSNs)における制限的なAPIによるデータアクセスの制限に起因する課題に対処すること。
- FacebookのAPI制限を回避する合法的なブラウザベースのデータ収集手法を開発すること。
- Facebookプロフィールから2つの主要なデータカテゴリである個人情報およびウォール活動を抽出・分析すること。
- 情報開示行動のパターン、特に性別による差異を特定すること。
- OSNデータ抽出および行動研究のためのWebクローラーの詳細な技術的設計を提供すること。
提案手法
- IMcrawlerはiMacros技術を用いて開発され、Facebookプロフィールからのブラウザベースのナビゲーションとデータ抽出を自動化する。
- クローラーは主に2つのデータタイプを収集する:個人情報(例:性別、場所、教育歴)およびウォール活動(例:投稿、いいね、コメント)。
- データ抽出は、自動ログインや公開可視性を超えるデータスクレイピングを回避することで、Facebookの利用規約を尊重する合法的な枠組み内で実施される。
- 収集されたデータは、統計的および意味的分析に適した2つの構造化されたデータセットに前処理される。
- 統計的分析が適用され、情報開示のパターン、活動頻度、属性間の関係が特定される。
- 個人属性と投稿内容および相互作用行動の相関関係を用いて、行動分析が可能になる。
実験結果
リサーチクエスチョン
- RQ1Facebookユーザーは、プロフィールで最も一般的にどのような個人情報を開示しているか?
- RQ2性別による違いは、Facebookプロフィールに明らかにされる情報のパターンや種別にどのように影響を及けるか?
- RQ3ユーザーのウォールに最も頻繁に投稿されるコンテンツの種別は何か?
- RQ4Facebookで最も一般的に見られるユーザー活動(例:投稿、いいね、コメント)は何か?
- RQ5個人属性とウォール活動の性質との間にどのような関係が存在するか?
主な発見
- 研究では、ユーザーが頻繁に性別、場所、教育背景といった個人属性をプロフィールに開示していることが明らかになった。
- 性別による情報開示のパターンの違いが観察され、コンテンツおよび属性共有における明確な傾向の差が確認された。
- ウォール活動の分析から、ステータス更新および写真投稿がFacebookで最も頻繁に行われる行動の一部であることが示された。
- 統計的分析により、個人属性とユーザーのウォールに共有されるコンテンツの種別との間に有意な相関関係が特定された。
- データ収集手法は、Facebookの利用規約に違反することなく包括的なデータセットを効果的に抽出できた。
- 提案されたクローラー設計は、再現可能で合法的なOSNデータ抽出フレームワークを提供し、行動研究において特に有用である。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。