The University of Osaka · 심리학
Changzeng Fu 교수의 연구실은 사회적 로봇이 인간과 장기적인 유대감을 형성할 수 있도록 하는 인공지능 기반 대화 시스템과 정서 인식 기술을 핵심으로 연구합니다. 특히 경험 기반 대화, 정서 공감 대화, 다중 모odal 정서 인식, 다국어 정서 인식 등 정서적 상호작용을 강화하는 기술 개발에 집중하고 있으며, 로봇이 인간의 정서를 이해하고 공감하는 데 기여하는 기반 기술을 구축하고자 합니다. 또한, 정서 기반 음성 변환 및 분류 모델을 통해 로봇의 정서적 표현 능력을 향상시키는 연구도 진행 중입니다.
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
Abstract Many social robots have emerged in public places to serve people. For these services, the robots are assumed to be able to present internal aspects (i.e., mind, sociability) to engage and interact with people over the long term. In this paper, we propose a novel dialogue structure called experience-based dialogue to help a robot present and maintain a good interaction over the long term. This dialogue structure contains a piece of knowledge and a story about how the robot gained this kn
Social connectedness is vital for developing group cohesion and strengthening belongingness. However, with the accelerating pace of modern life, people have fewer opportunities to participate in group-building activities. Furthermore, owing to the teleworking and quarantine requirements necessitated by the Covid-19 pandemic, the social connectedness of group members may become weak. To address this issue, in this study, we used an android robot to conduct daily conversations, and as an intermedi
Mental health issues are receiving more and more attention in society. In this paper, we introduce a preliminary study on human-robot mental comforting conversation, to make an android robot (ERICA) present an understanding of the user's situation by sharing similar emotional experiences to enhance the perception of empathy. Specifically, we create the emotional speech for ERICA by using CycleGAN-based emotional voice conversion model, in which the pitch and spectrogram of the speech are convert
In this study, we try to recognize the similarities between different languages in expressing basic human emotions by cross/multi-language corpus training of a novel recognition model based on one dimesional convolutional neural network (CNN) and bi-directional long short-term memory (bi-LSTM) with attention mechanism, we named it CAbiLS. We train and test the model using various combinations of three different corpora of three different languages (German, Chinese and Italian) and we also discus
Speaker individual bias may cause emotion-related features to form clusters with irregular borders (non-Gaussian distributions), making the model sensitive to local irregularities of pattern distributions, resulting in the model over-fit of the in-domain dataset. This problem may cause a decrease in the validation scores in cross-domain (i.e., speaker-independent, channel-variant) implementation. To mitigate this problem, in this paper, we propose an adversarial training-based classifier to regu
Emotion recognition has been gaining attention in recent years due to its applications on artificial agents. To achieve a good performance with this task, much research has been conducted on the multi-modality emotion recognition model for leveraging the different strengths of each modality. However, a research question remains: what exactly is the most appropriate way to fuse the information from different modalities? In this paper, we proposed audio sample augmentation and an emotion-oriented
In recent years, automatic emotion recognition has attracted the attention of researchers because of its great effects and wide implementations in supporting humans' activities. Given that the data about emotions is difficult to collect and organize into a large database like the dataset of text or images, the true distribution would be difficult to be completely covered by the training set, which affects the model's robustness and generalization in subsequent applications. In this paper, we pro
In this paper, we propose an adversarial auto-encoder-based classifier, which can regularize the distribution of latent representation to smooth the boundaries among categories. Moreover, we adopt multi-instance learning by dividing speech into a bag of segments to capture the most salient moments for presenting an emotion. The proposed model was trained on the IEMOCAP dataset and evaluated on the in-corpus validation set (IEMOCAP) and the cross-corpus validation set (MELD). The experiment resul
In this paper, we propose an attention-based CNN-BLSTM model with the end-to-end (E2E) learning method. We first extract Mel-spectrogram from wav file instead of using handcrafted features. Then we adopt two types of attention mechanisms to let the model focuses on salient periods of speech emotions over the temporal dimension. Considering that there are many individual differences among people in expressing emotions, we incorporate speaker recognition as an auxiliary task. Moreover, since the t
In this study, we explore the transformer's ability to capture intra-relations among frames by augmenting the receptive field of models. Concretely, we propose a CycleGAN-based model with the transformer and investigate its ability in the emotional voice conversion task. In the training procedure, we adopt curriculum learning to gradually increase the frame length so that the model can see from the short segment till the entire speech. The proposed method was evaluated on the Japanese emotional