Yonsei University · 情報科学
Professor Min Song's research lab specializes in text and web mining, with a focus on analyzing public discourse on health issues, scientific knowledge discovery, and social media dynamics. The lab develops advanced computational methods—such as topic modeling, entity-network analysis, and sentiment scoring—to extract meaningful insights from large-scale textual data from sources like Twitter, news articles, and scientific literature. A key emphasis is on understanding how individuals share health experiences online and how knowledge units (e.g., drugs, biological entities) are interconnected and cited in scientific literature. The lab also pioneers novel metrics like 'entitymetrics' to quantify the impact of knowledge entities in scientific networks.
Figures are computed from collected data and may differ slightly.
The present study investigates topic coverage and sentiment dynamics of two different media sources, Twitter and news publications, on the hot health issue of Ebola. We conduct content and sentiment analysis by: (1) applying vocabulary control to collected datasets; (2) employing the n-gram LDA topic modeling technique; (3) adopting entity extraction and entity network; and (4) introducing the concept of topic-based sentiment scores. With the query term ‘Ebola’ or ‘Ebola virus’, we collected 16,
This paper proposes entitymetrics to measure the impact of knowledge units. Entitymetrics highlight the importance of entities embedded in scientific literature for further knowledge discovery. In this paper, we use Metformin, a drug for diabetes, as an example to form an entity-entity citation network based on literature related to Metformin. We then calculate the network features and compare the centrality ranks of biological entities with results from Comparative Toxicogenomics Database (CTD)
"This handbook presents recent advances and surveys of applications in text and web mining of interests to researchers and endusers " - Provided by publisher
This study showed that people are eager to share their personal experience with chronic diseases on social media platforms despite possible privacy and security issues. The results reported in this paper are promising and demonstrate the need for more in-depth studies on the way patients with chronic diseases express themselves on social media platforms.
The results imply that the performance of dictionary-based extraction techniques is largely influenced by information resources used to build the dictionary. In addition, the edit distance algorithm shows steady performance with three different dictionaries in precision whereas the context-only technique achieves a high-end performance with three difference dictionaries in recall.
Social media is changing existing information behavior by giving users access to real-time online information channels without the constraints of time and space. Social media, therefore, has created an enormous data analysis challenge for scientists trying to keep pace with developments in their field. Most previous studies have adopted broad-brush approaches that typically result in limited analysis possibilities. To address this problem, we applied text-mining techniques to Twitter data relate
Open papers in the app to read, cite, and organize with AI.