Min Song
연세대학교 컴퓨터학과 · 컴퓨터과학
Min Song 교수의 연구실은 텍스트 및 웹 마이닝 기반의 지식 탐색과 정보 분석을 핵심으로 하며, 특히 소셜 미디어와 과학문헌에서의 엔티티 기반 지식 추출, 감성 동적 분석, 주제 모델링 기법을 활용한 실시간 정보 추적에 집중하고 있습니다. 의료 이슈나 사회적 이슈에 대한 언론과 SNS의 언어적 표현을 비교 분석함으로써, 정보의 흐름과 공중 여론의 변화를 정량적으로 해석하는 데 기여하고 있습니다. 또한, 과학문헌 내 생물학적 엔티티 간의 상호작용 네트워크를 분석해 지식 유출 가능성을 평가하는 '엔티티미터틱스(Entitymetrics)'와 같은 혁신적 분석 프레임워크를 개발하고 있습니다.
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
The present study investigates topic coverage and sentiment dynamics of two different media sources, Twitter and news publications, on the hot health issue of Ebola. We conduct content and sentiment analysis by: (1) applying vocabulary control to collected datasets; (2) employing the n-gram LDA topic modeling technique; (3) adopting entity extraction and entity network; and (4) introducing the concept of topic-based sentiment scores. With the query term ‘Ebola’ or ‘Ebola virus’, we collected 16,
This paper proposes entitymetrics to measure the impact of knowledge units. Entitymetrics highlight the importance of entities embedded in scientific literature for further knowledge discovery. In this paper, we use Metformin, a drug for diabetes, as an example to form an entity-entity citation network based on literature related to Metformin. We then calculate the network features and compare the centrality ranks of biological entities with results from Comparative Toxicogenomics Database (CTD)
"This handbook presents recent advances and surveys of applications in text and web mining of interests to researchers and endusers " - Provided by publisher
This study showed that people are eager to share their personal experience with chronic diseases on social media platforms despite possible privacy and security issues. The results reported in this paper are promising and demonstrate the need for more in-depth studies on the way patients with chronic diseases express themselves on social media platforms.
The results imply that the performance of dictionary-based extraction techniques is largely influenced by information resources used to build the dictionary. In addition, the edit distance algorithm shows steady performance with three different dictionaries in precision whereas the context-only technique achieves a high-end performance with three difference dictionaries in recall.
Social media is changing existing information behavior by giving users access to real-time online information channels without the constraints of time and space. Social media, therefore, has created an enormous data analysis challenge for scientists trying to keep pace with developments in their field. Most previous studies have adopted broad-brush approaches that typically result in limited analysis possibilities. To address this problem, we applied text-mining techniques to Twitter data relate