[논문 리뷰] What Twitter Data Tell Us about the Future?
이 연구는 트위터 퓨처리스트의 100만 건 이상의 트윗을 LDA와 BERTopic를 사용해 분석하여 예측 가능한 미래를 모델링하며, LDA를 통해 15개의 주제를, BERTopic를 통해 100개의 고유한 주제를 파악했다. 연구는 '미래 현재(future present)' 개념—동적인, 진화하는 잠재적 미래의 개념적 묘사—를 규명하여, 이들이 사회적 미디어 사용자가 미래를 예측하고 사전에 대응하도록 이끄는 방식을 보여주며, 영향력 있는 목소리의 언어적 신호가 확장 가능한 NLP 파이프라인과 오픈소스로 공개된 데이터 및 코드를 통해 집단적 예측 행동을 형성함을 입증한다.
Anticipation is a fundamental human cognitive ability that involves thinking about and living towards the future. While language markers reflect anticipatory thinking, research on anticipation from the perspective of natural language processing is limited. This study aims to investigate the futures projected by futurists on Twitter and explore the impact of language cues on anticipatory thinking among social media users. We address the research questions of what futures Twitter's futurists anticipate and share, and how these anticipated futures can be modeled from social data. To investigate this, we review related works on anticipation, discuss the influence of language markers and prestigious individuals on anticipatory thinking, and present a taxonomy system categorizing futures into "present futures" and "future present". This research presents a compiled dataset of over 1 million publicly shared tweets by future influencers and develops a scalable NLP pipeline using SOTA models. The study identifies 15 topics from the LDA approach and 100 distinct topics from the BERTopic approach within the futurists' tweets. These findings contribute to the research on topic modelling and provide insights into the futures anticipated by Twitter's futurists. The research demonstrates the futurists' language cues signals futures-in-the-making that enhance social media users to anticipate their own scenarios and respond to them in present. The fully open-sourced dataset, interactive analysis, and reproducible source code are available for further exploration.
연구 동기 및 목표
- 트위터 기반 퓨처리스트가 예측하고 공유하는 미래의 유형과 그것이 대중의 예측적 사고에 미치는 영향을 조사하는 것.
- 언어적 표시어와 유명 인물의 역할이 소셜 미디어 사용자의 미래 시나리오에 대한 예측 행동을 어떻게 형성하는지 분석하는 것.
- 소셜 미디어 데이터에서 예측 가능한 미래를 모델링하기 위한 확장 가능한 NLP 파이프라인 개발.
- 미래 연구를 위한 예측 어조와 사회적 예측에 관한 향후 연구를 위해 공개 가능하고 재현 가능한 데이터셋 및 분석 프레임워크 구축.
제안 방법
- 트위터 학술 API의 제약과 데이터 가용성 고려사항을 충족시키기 위해 웹 스크래핑을 통해 공개된 트위터 퓨처리스트의 100만 건 이상의 트윗을 수집.
- 텍스트 데이터를 정제, 토큰화, 노이즈 제거 등의 표준 NLP 기법을 통해 전처리하여 주제 모델링에 대비.
- LDA(Latent Dirichlet Allocation)를 적용해 퓨처리스트 트윗 내 15개의 광범위한 주제 군집을 식별.
- Transformer 기반 주제 모델링 기법인 BERTopic을 활용해 동일한 데이터셋에서 100개의 더 세밀하고 의미적으로 일관된 주제를 추출.
- 맥락 기반 임bedding과 클러스터링을 활용해 짧은 소셜 미디어 텍스트에 특화된 확장 가능한 최신 기술 기반(NLP) 파이프라인 개발.
- 모든 데이터셋, 상호작용 가능한 시각화 자료, 소스 코드를 오픈 사이언스 프레임워크(Open Science Framework)를 통해 공개하여 재현성과 커뮤니티 재사용을 보장.

실험 결과
연구 질문
- RQ1RQ.1: 트위터의 퓨처리스트는 어떤 미래를 예측하고 공유하는가?
- RQ2RQ.2: NLP 기법을 사용해 소셜 데이터에서 예측 가능한 미래를 어떻게 모델링할 수 있는가?
- RQ3RQ.3: 영향력 있는 퓨처리스트의 언어적 신호는 소셜 미디어 사용자의 예측 행동에 어떤 영향을 미치는가?
- RQ4RQ.4: 온라인 논의에서 '미래 현재' 개념의 구조와 진화는 어떠한가?
주요 결과
- LDA 모델은 트위터 퓨처리스트 트윗 내 15개의 광범위한 주제 군집을 식별하여 고수준의 예측적 서사 구조를 반영했다.
- BERTopic 모델은 100개의 고유하고 의미적으로 풍부한 주제를 파악하여 LDA 대비 더 높은 세분화 수준과 해석 가능성의 우수성을 입증했다.
- 다수의 예측된 미래는 '미래 현재(future present)'로 분류되었는데, 이는 동적인, 진화하는, 예측 불가능한 잠재적 미래의 개념적 묘사였다.
- 이러한 '미래 현재' 미래는 고정된 것이 아니라 살아있는, 유연하게 변형 가능한 비전을 나타내며, 사용자가 사전에 예측하고 대응하도록 이끈다.
- 유명한 퓨처리스트의 언어적 신호는 소셜 미디어 사용자가 다양한 가능성을 가진 미래를 예측하고 준비하는 데 능력을 향상시키는 신호로 작용한다.
- 본 연구는 소셜 미디어 데이터, 특히 사고 리더의 데이터가 체계적으로 예측적 서사를 추출하고 분석할 수 있음을 입증한다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.