Skip to main content
QUICK REVIEW

[Paper Review] Improving the quality of individual-level online information tracking: challenges of existing approaches and introduction of a new content- and long-tail sensitive academic solution

Silke Adam, Mykola Makhortykh|arXiv (Cornell University)|Mar 5, 2024
Web visibility and informetrics4 citations
TL;DR

This paper introduces WebTrack, an open-source, content- and long-tail-sensitive academic tool for individual-level online information tracking that overcomes limitations in existing desktop tracking tools by capturing detailed content exposure beyond news lists. Using data from 1,185 participants, WebTrack enables more accurate detection of politics-related information consumption through automated content analysis, significantly improving data quality and enabling novel analytical insights in social science research.

ABSTRACT

This article evaluates the quality of data collection in individual-level desktop information tracking used in the social sciences and shows that the existing approaches face sampling issues, validity issues due to the lack of content-level data and their disregard of the variety of devices and long-tail consumption patterns as well as transparency and privacy issues. To overcome some of these problems, the article introduces a new academic tracking solution, WebTrack, an open source tracking tool maintained by a major European research institution. The design logic, the interfaces and the backend requirements for WebTrack, followed by a detailed examination of strengths and weaknesses of the tool, are discussed. Finally, using data from 1185 participants, the article empirically illustrates how an improvement in the data collection through WebTrack leads to new innovative shifts in the processing of tracking data. As WebTrack allows collecting the content people are exposed to on more than classical news platforms, we can strongly improve the detection of politics-related information consumption in tracking data with the application of automated content analysis compared to traditional approaches that rely on the list-based identification of news.

Motivation & Objective

  • To identify critical flaws in existing desktop-based individual-level online information tracking tools, particularly regarding sampling bias, lack of content-level data, and disregard for long-tail content and device diversity.
  • To address validity, transparency, and privacy issues in current tracking methodologies used in social science research.
  • To develop and evaluate a new academic tracking solution—WebTrack—that captures detailed content exposure across diverse online platforms.
  • To demonstrate how enhanced data collection through WebTrack enables more accurate and innovative processing of tracking data, especially for detecting politics-related information consumption.

Proposed method

  • Designing WebTrack as a desktop-based tracking tool with a client-side browser extension and a secure server-side backend for data aggregation and storage.
  • Implementing content-level logging that captures full URLs, page titles, and raw HTML content to enable automated content analysis.
  • Integrating device- and platform-agnostic tracking to account for long-tail content consumption across non-traditional news sources.
  • Applying automated content analysis techniques to classify tracked content by topic, particularly politics, using NLP-based classification models.
  • Ensuring privacy and transparency through anonymization, user consent mechanisms, and open-source code availability.
  • Validating the tool’s performance and data quality using a large-scale empirical study with 1,185 participants across diverse online behaviors.

Experimental results

Research questions

  • RQ1How do existing individual-level online tracking tools fail in capturing content diversity and long-tail consumption patterns?
  • RQ2To what extent does WebTrack improve the accuracy of detecting politics-related information exposure compared to list-based tracking methods?
  • RQ3What are the key technical and ethical challenges in deploying a content-sensitive, long-tail-aware tracking system in academic research?
  • RQ4How does the inclusion of detailed content data enable new analytical capabilities in processing online information exposure?

Key findings

  • WebTrack successfully captures content exposure across a broad range of platforms, including non-traditional news sources, significantly expanding the scope of detectable information consumption beyond classical news lists.
  • The integration of automated content analysis with WebTrack enables a more precise identification of politics-related content, reducing reliance on potentially inaccurate list-based classification.
  • Empirical data from 1,185 participants demonstrate that WebTrack collects richer, more representative data on online media exposure, especially for niche and long-tail content.
  • The tool improves data validity by accounting for device diversity and user-specific browsing behavior, reducing sampling bias inherent in traditional tracking approaches.
  • WebTrack enhances transparency and privacy through open-source design, user consent workflows, and secure data handling, addressing key ethical concerns in tracking research.
  • The enhanced data quality enables novel analytical shifts, such as identifying subtle political content exposure patterns previously undetectable with conventional methods.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.