Alice Oh
Korea Advanced Institute of Science and Technology · 情報科学
研究室紹介
Professor Alice Oh's research lab specializes in natural language processing, multimodal affective computing, and graph-based machine learning, with a focus on understanding human behavior and sentiment in real-world contexts. The lab develops advanced models for aspect-based sentiment analysis, emotion recognition in natural social interactions, and robust representation learning in noisy graph structures. Their work emphasizes scalable, self-supervised, and human-centered approaches to extract meaningful insights from unstructured and multimodal data such as online reviews, social media, and sensor-based affective signals. The lab bridges theoretical modeling with practical applications in misinformation detection, dialogue systems, and mental health monitoring.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15User-generated reviews on the Web contain sentiments about detailed aspects of products and services. However, most of the reviews are plain text and thus require much effort to obtain information about relevant details. In this paper, we tackle the problem of automatically discovering what aspects are evaluated in reviews and how sentiments for different aspects are expressed. We first propose Sentence-LDA (SLDA), a probabilistic generative model that assumes all words in a single sentence are
Online social networking sites are experimenting with the following crowd-powered procedure to reduce the spread of fake news and misinformation: whenever a user is exposed to a story through her feed, she can flag the story as misinformation and, if the story receives enough flags, it is sent to a trusted third party for fact checking. If this party identifies the story as misinformation, it is marked as disputed. However, given the uncertain number of exposures, the high cost of fact checking,
The two current approaches to language generation, template-based and rule-based (linguistic) NLG, have limitations when applied to spoken dialogue systems, in part because they were developed for text generation. In this paper, we propose a new corpus-based approach to natural language generation, specifically designed for spoken dialogue systems.
ABSTRACT: Recognizing emotions during social interactions has many potential applications with the popularization of low-cost mobile sensors, but a challenge remains with the lack of naturalistic affective interaction data. Most existing emotion datasets do not support studying idiosyncratic emotions arising in the wild as they were collected in constrained environments. Therefore, studying emotions in the context of social interactions requires a novel dataset, and K-EmoCon is such a multimodal
To help users quickly understand the major opinions from massive online reviews, it is important to automatically reveal the latent structure of the aspects, sentiment polarities, and the association between them. However, there is little work available to do this effectively. In this paper, we propose a hierarchical aspect sentiment model (HASM) to discover a hierarchical structure of aspect-based sentiments from unlabeled online reviews. In HASM, the whole structure is a tree. Each node itself
Attention mechanism in graph neural networks is designed to assign larger weights to important neighbor nodes for better representation. However, what graph attention learns is not understood well, particularly when graphs are noisy. In this paper, we propose a self-supervised graph attention network (SuperGAT), an improved graph attention model for noisy graphs. Specifically, we exploit two attention forms compatible with a self-supervised task to predict edges, whose presence and absence conta
We present a computational framework for understanding the social aspects of emotions in Twitter conversations. Using unannotated data and semisupervised machine learning, we look at emotional transitions, emotional influences among the conversation partners, and patterns in the overall emotional exchanges. We find that conversational partners usually express the same emotion, which we name Emotion accommodation, but when they do not, one of the conversational partners tends to respond with a po
We discuss our findings from a study using Twitter lists to infer the characteristics and interests of users. Gathering and structuring user interest has been challenging because it often requires expensive and/or proprietary data such as users' clickthrough logs or desktop histories. We show that by using the tweets of all the users in a Twitter list, we can discover characteristics and interests of the users in that list, even if the users as individuals do not tweet about those interests. We
We present "look-to-talk", a gaze-aware interface for directing a spoken utterance to a software agent in a multi-user collaborative environment. Through a prototype and a Wizard-of-Oz (Woz) experiment, we show that "look-to-talk" is indeed a natural alternative to speech and other paradigms.
Multilingualism is common offline, but we have a more limited understanding of the ways multilingualism is displayed online and the roles that multilinguals play in the spread of content between speakers of different languages. We take a computational approach to studying multilingualism using one of the largest user-generated content platforms, Wikipedia. We study multilingualism by collecting and analyzing a large dataset of the content written by multilingual editors of the English, German, a
This paper presents a Web-based user evaluation of a system for classifying and presenting political viewpoints of blog posts. The system is based on a classification model trained using a supervised learning algorithm, and the data set consists of recent posts from blogs that are self-identified as a liberal or a conservative viewpoint. We first discuss the classification process. Then, with a prototype system for retrieving and classifying political blogs, we look at how the classification res
The two current approaches to language generation, template-based and rule-based (linguistic) NLG, have limitations when applied to spoken dialogue systems, in part because they were developed for text generation. In this paper, we propose a new corpus-based approach to natural language generation, specifically designed for spoken dialogue systems.
this paper, we describe our MeetingManager system, a multiuser multimodal collaboration tool for planning, facilitating, and browsing structured meetings