The University of Osaka · Psychology
Professor Changzeng Fu's research lab specializes in human-robot interaction, with a focus on enhancing social robots' ability to foster long-term engagement and emotional connection. The lab explores experience-based dialogue systems, empathetic robot communication, and multimodal emotion recognition to enable robots to perceive, express, and respond to human emotions naturally. Key research directions include affective computing, speech emotion synthesis, and cross-lingual emotion recognition using deep learning models such as CNN-BiLSTM with attention and graph-based fusion techniques.
Figures are computed from collected data and may differ slightly.
Abstract Many social robots have emerged in public places to serve people. For these services, the robots are assumed to be able to present internal aspects (i.e., mind, sociability) to engage and interact with people over the long term. In this paper, we propose a novel dialogue structure called experience-based dialogue to help a robot present and maintain a good interaction over the long term. This dialogue structure contains a piece of knowledge and a story about how the robot gained this kn
Social connectedness is vital for developing group cohesion and strengthening belongingness. However, with the accelerating pace of modern life, people have fewer opportunities to participate in group-building activities. Furthermore, owing to the teleworking and quarantine requirements necessitated by the Covid-19 pandemic, the social connectedness of group members may become weak. To address this issue, in this study, we used an android robot to conduct daily conversations, and as an intermedi
Mental health issues are receiving more and more attention in society. In this paper, we introduce a preliminary study on human-robot mental comforting conversation, to make an android robot (ERICA) present an understanding of the user's situation by sharing similar emotional experiences to enhance the perception of empathy. Specifically, we create the emotional speech for ERICA by using CycleGAN-based emotional voice conversion model, in which the pitch and spectrogram of the speech are convert
In this study, we try to recognize the similarities between different languages in expressing basic human emotions by cross/multi-language corpus training of a novel recognition model based on one dimesional convolutional neural network (CNN) and bi-directional long short-term memory (bi-LSTM) with attention mechanism, we named it CAbiLS. We train and test the model using various combinations of three different corpora of three different languages (German, Chinese and Italian) and we also discus
Speaker individual bias may cause emotion-related features to form clusters with irregular borders (non-Gaussian distributions), making the model sensitive to local irregularities of pattern distributions, resulting in the model over-fit of the in-domain dataset. This problem may cause a decrease in the validation scores in cross-domain (i.e., speaker-independent, channel-variant) implementation. To mitigate this problem, in this paper, we propose an adversarial training-based classifier to regu
Emotion recognition has been gaining attention in recent years due to its applications on artificial agents. To achieve a good performance with this task, much research has been conducted on the multi-modality emotion recognition model for leveraging the different strengths of each modality. However, a research question remains: what exactly is the most appropriate way to fuse the information from different modalities? In this paper, we proposed audio sample augmentation and an emotion-oriented
In recent years, automatic emotion recognition has attracted the attention of researchers because of its great effects and wide implementations in supporting humans' activities. Given that the data about emotions is difficult to collect and organize into a large database like the dataset of text or images, the true distribution would be difficult to be completely covered by the training set, which affects the model's robustness and generalization in subsequent applications. In this paper, we pro
In this paper, we propose an adversarial auto-encoder-based classifier, which can regularize the distribution of latent representation to smooth the boundaries among categories. Moreover, we adopt multi-instance learning by dividing speech into a bag of segments to capture the most salient moments for presenting an emotion. The proposed model was trained on the IEMOCAP dataset and evaluated on the in-corpus validation set (IEMOCAP) and the cross-corpus validation set (MELD). The experiment resul
In this paper, we propose an attention-based CNN-BLSTM model with the end-to-end (E2E) learning method. We first extract Mel-spectrogram from wav file instead of using handcrafted features. Then we adopt two types of attention mechanisms to let the model focuses on salient periods of speech emotions over the temporal dimension. Considering that there are many individual differences among people in expressing emotions, we incorporate speaker recognition as an auxiliary task. Moreover, since the t
In this study, we explore the transformer's ability to capture intra-relations among frames by augmenting the receptive field of models. Concretely, we propose a CycleGAN-based model with the transformer and investigate its ability in the emotional voice conversion task. In the training procedure, we adopt curriculum learning to gradually increase the frame length so that the model can see from the short segment till the entire speech. The proposed method was evaluated on the Japanese emotional
Open papers in the app to read, cite, and organize with AI.