[Paper Review] Community Question Answering Platforms vs. Twitter for Predicting Characteristics of Urban Neighbourhoods
This paper investigates whether text from Yahoo! Answers, a community question answering (QA) platform, can predict real-world demographic attributes of London neighborhoods, comparing it to Twitter microblogs. Using NLP and regression models on user-generated text, it finds that QA text predicts demographics with a mean Pearson correlation of ρ = 0.54, slightly outperforming Twitter (ρ = 0.53), and reveals distinct semantic patterns between the two platforms: QA offers encyclopedic knowledge, while Twitter reflects current sociocultural trends.
In this paper, we investigate whether text from a Community Question Answering (QA) platform can be used to predict and describe real-world attributes. We experiment with predicting a wide range of 62 demographic attributes for neighbourhoods of London. We use the text from QA platform of Yahoo! Answers and compare our results to the ones obtained from Twitter microblogs. Outcomes show that the correlation between the predicted demographic attributes using text from Yahoo! Answers discussions and the observed demographic attributes can reach an average Pearson correlation coefficient of {ho} = 0.54, slightly higher than the predictions obtained using Twitter data. Our qualitative analysis indicates that there is semantic relatedness between the highest correlated terms extracted from both datasets and their relative demographic attributes. Furthermore, the correlations highlight the different natures of the information contained in Yahoo! Answers and Twitter. While the former seems to offer a more encyclopedic content, the latter provides information related to the current sociocultural aspects or phenomena.
Motivation & Objective
- To evaluate whether non-geotagged, community-generated text from QA platforms like Yahoo! Answers can predict real-world demographic attributes of urban neighborhoods.
- To compare the predictive performance of Yahoo! Answers text against Twitter microblogs for 62 demographic attributes of London neighborhoods.
- To analyze the semantic relationship between high-coefficient terms in text and their corresponding demographic attributes.
- To contrast the nature of information in QA platforms (encyclopedic) versus social media (current sociocultural) in urban attribute modeling.
- To demonstrate that QA text, despite lacking geolocation, can yield strong predictive signals for urban planning and sociological analysis.
Proposed method
- Text from Yahoo! Answers discussions about London neighborhoods was collected and preprocessed using standard NLP techniques, including tokenization, stopword removal, and lemmatization.
- A logistic regression model was trained to predict 62 demographic attributes using TF-IDF weighted features extracted from QA and Twitter text.
- Model performance was evaluated using 10-fold cross-validation, with Pearson correlation coefficient (ρ) as the primary metric.
- Terms with the highest absolute coefficients in the predictive models were extracted to analyze semantic relevance to demographic attributes.
- Qualitative analysis compared the most correlated terms from both platforms to their associated demographic attributes, highlighting differences in content type.
- A Wilcoxon signed-rank test was used to assess statistical significance of performance differences between Yahoo! Answers and Twitter.
Experimental results
Research questions
- RQ1Can text from a community question answering platform like Yahoo! Answers predict real-world demographic attributes of urban neighborhoods with comparable or better performance than Twitter?
- RQ2How do the semantic patterns of high-impact terms in Yahoo! Answers and Twitter text relate to the demographic attributes they predict?
- RQ3What are the differences in the nature of information captured by QA platforms versus microblogging platforms in describing urban neighborhoods?
- RQ4Does the lack of geolocation in QA data hinder its predictive power for neighborhood-level attributes?
- RQ5Are certain demographic attributes (e.g., ethnicity, employment, age group) better predicted by one platform over the other?
Key findings
- The average Pearson correlation coefficient (ρ) for predicting 62 demographic attributes using Yahoo! Answers text was 0.54, slightly higher than Twitter’s 0.53.
- Yahoo! Answers outperformed Twitter in predicting attributes related to ethnicity and employment, while Twitter performed better for age group and car ownership.
- The Wilcoxon signed-rank test confirmed the performance difference between the two platforms was statistically significant (p < 0.01).
- Qualitative analysis revealed that the most predictive terms from both platforms showed strong semantic alignment with their corresponding demographic attributes, such as 'Jewish' and 'Jewish communities' in discussions about religious demographics.
- Text from Yahoo! Answers was characterized as more encyclopedic and factual, whereas Twitter text reflected real-time sociocultural sentiments and current events.
- Despite the absence of geolocation data in Yahoo! Answers, the platform’s text still yielded robust predictive signals for neighborhood-level demographics.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.