Skip to main content
QUICK REVIEW

[논문 리뷰] Did you take the pill? - Detecting Personal Intake of Medicine from Twitter

Debanjan Mahata, Jasper Friedrichs|arXiv (Cornell University)|2018. 08. 03.
Spam and Phishing Detection참고 문헌 19인용 수 6
한 줄 요약

이 논문은 수동으로 애너테이션된 데이터셋을 사용하여 트위터 게시물에서 개인 약물 복용 여부를 탐지하기 위해 얕은 합성곱 신경망(CNN)의 스태킹 앙상블을 제안한다. 이 시스템은 최신 기술 수준의 마이크로 평균 F-스코어 0.693을 달성하여 약물 감시, 정신 건강 모니터링 및 정서 컴퓨팅 분야에서 개인 수준의 약물 복용 추적을 가능하게 한다.

ABSTRACT

Mining social media messages such as tweets, articles, and Facebook posts for health and drug related information has received significant interest in pharmacovigilance research. Social media sites (e.g., Twitter), have been used for monitoring drug abuse, adverse reactions of drug usage and analyzing expression of sentiments related to drugs. Most of these studies are based on aggregated results from a large population rather than specific sets of individuals. In order to conduct studies at an individual level or specific cohorts, identifying posts mentioning intake of medicine by the user is necessary. Towards this objective we develop a classifier for identifying mentions of personal intake of medicine in tweets. We train a stacked ensemble of shallow convolutional neural network (CNN) models on an annotated dataset. We use random search for tuning the hyper-parameters of the CNN models and present an ensemble of best models for the prediction task. Our system produces state-of-the-art result, with a micro-averaged F-score of 0.693. We believe that the developed classifier has direct uses in the areas of psychology, health informatics, pharmacovigilance and affective computing for tracking moods, emotions and sentiments of patients expressing intake of medicine in social media.

연구 동기 및 목표

  • 소셜 미디어에서 개인 약물 복용을 식별하여 개인 수준의 건강 모니터링을 가능하게 하기 위해.
  • 기존의 약물 감시 연구가 집단 수준의 추세에만 초점을 맞추는 데 반해 개인 행동을 다루지 못하는 격차를 메우기 위해.
  • 트위터와 같은 짧고 비공식적인 텍스트에서 자기 보고된 약물 복용을 탐지하기 위한 강력한 분류기 개발을 위해.
  • 정서 컴퓨팅 및 건강 정보학 분야의 애플리케이션을 지원하기 위해 약물 복용과 관련된 기분 및 감정 추적을 위해.

제안 방법

  • 저자는 트위터 게시물의 수동 애너테이션된 데이터셋을 기반으로 얕은 합성곱 신경망(CNN) 모델의 스태킹 앙상블을 훈련시킨다.
  • CNN 모델의 하이퍼파rameter는 성능 최적화를 위해 무작위 탐색을 사용하여 조정된다.
  • 최종 예측은 앙상블에서 성능이 가장 뛰어난 모델들의 출력을 조합하여 이루어진다.
  • 데이터셋은 일반적인 약물 언급과 개인 복용 여부를 구분하기 위해 수동 애너테이션을 통해 구성된다.
  • 모델는 주로 마이크로 평균 F-스코어를 주요 평가 지표로 사용하여 훈련 및 평가된다.

실험 결과

연구 질문

  • RQ1딥 러닝 모델이 비공식적인 소셜 미디어 텍스트에서 개인 약물 복용 여부를 효과적으로 탐지할 수 있는가?
  • RQ2CNN 모델의 앙상블가 단일 모델보다 자기 보고된 약물 복용을 식별하는 데 더 나은 성능을 보이는가?
  • RQ3개인 약물 복용 탐지의 벤치마크 데이터셋에서 달성할 수 있는 성능 수준은 어느 정도인가?
  • RQ4이 분류기는 정신 건강 및 약물 감시 분야에서 개인 수준의 건강 모니터링을 어느 정도 지원할 수 있는가?

주요 결과

  • 제안된 CNN의 스태킹 앙상블은 마이크로 평균 F-스코어 0.693을 달성하여 이 작업 분야에서 최신 기술 수준의 결과를 보여준다.
  • 하이퍼파rameter 조정을 위해 무작위 탐색을 사용함으로써 모델의 일반화 능력과 성능이 향상되었다.
  • 분류기는 트위터에서 일반적 언급이나 제3자 언급과 개인 복용 언급을 성공적으로 구분한다.
  • 모델는 정신 건강 및 정서 컴퓨팅 응용 분야에서 환자 보고 약물 복용 추적에 실질적인 유용성을 보여준다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.