大阪大学 · 心理学
福昌増教授の研究室は、人間とロボットの持続的で感情的な関係構築に焦点を当てた研究を進めています。特に、ロボットが過去の対話経験を共有することで人間との絆を強化する「経験基盤対話」や、感情認識に基づく対話的支援技術の開発が特徴です。音声の感情変換やマルチモodal感情認識、ドメイン一般化のための正則化手法など、実用的で社会的価値の高いロボット対話技術の基盤を構築しています。
Figures are computed from collected data and may differ slightly.
Abstract Many social robots have emerged in public places to serve people. For these services, the robots are assumed to be able to present internal aspects (i.e., mind, sociability) to engage and interact with people over the long term. In this paper, we propose a novel dialogue structure called experience-based dialogue to help a robot present and maintain a good interaction over the long term. This dialogue structure contains a piece of knowledge and a story about how the robot gained this kn
Social connectedness is vital for developing group cohesion and strengthening belongingness. However, with the accelerating pace of modern life, people have fewer opportunities to participate in group-building activities. Furthermore, owing to the teleworking and quarantine requirements necessitated by the Covid-19 pandemic, the social connectedness of group members may become weak. To address this issue, in this study, we used an android robot to conduct daily conversations, and as an intermedi
Mental health issues are receiving more and more attention in society. In this paper, we introduce a preliminary study on human-robot mental comforting conversation, to make an android robot (ERICA) present an understanding of the user's situation by sharing similar emotional experiences to enhance the perception of empathy. Specifically, we create the emotional speech for ERICA by using CycleGAN-based emotional voice conversion model, in which the pitch and spectrogram of the speech are convert
In this study, we try to recognize the similarities between different languages in expressing basic human emotions by cross/multi-language corpus training of a novel recognition model based on one dimesional convolutional neural network (CNN) and bi-directional long short-term memory (bi-LSTM) with attention mechanism, we named it CAbiLS. We train and test the model using various combinations of three different corpora of three different languages (German, Chinese and Italian) and we also discus
Speaker individual bias may cause emotion-related features to form clusters with irregular borders (non-Gaussian distributions), making the model sensitive to local irregularities of pattern distributions, resulting in the model over-fit of the in-domain dataset. This problem may cause a decrease in the validation scores in cross-domain (i.e., speaker-independent, channel-variant) implementation. To mitigate this problem, in this paper, we propose an adversarial training-based classifier to regu
Emotion recognition has been gaining attention in recent years due to its applications on artificial agents. To achieve a good performance with this task, much research has been conducted on the multi-modality emotion recognition model for leveraging the different strengths of each modality. However, a research question remains: what exactly is the most appropriate way to fuse the information from different modalities? In this paper, we proposed audio sample augmentation and an emotion-oriented
In recent years, automatic emotion recognition has attracted the attention of researchers because of its great effects and wide implementations in supporting humans' activities. Given that the data about emotions is difficult to collect and organize into a large database like the dataset of text or images, the true distribution would be difficult to be completely covered by the training set, which affects the model's robustness and generalization in subsequent applications. In this paper, we pro
In this paper, we propose an adversarial auto-encoder-based classifier, which can regularize the distribution of latent representation to smooth the boundaries among categories. Moreover, we adopt multi-instance learning by dividing speech into a bag of segments to capture the most salient moments for presenting an emotion. The proposed model was trained on the IEMOCAP dataset and evaluated on the in-corpus validation set (IEMOCAP) and the cross-corpus validation set (MELD). The experiment resul
In this paper, we propose an attention-based CNN-BLSTM model with the end-to-end (E2E) learning method. We first extract Mel-spectrogram from wav file instead of using handcrafted features. Then we adopt two types of attention mechanisms to let the model focuses on salient periods of speech emotions over the temporal dimension. Considering that there are many individual differences among people in expressing emotions, we incorporate speaker recognition as an auxiliary task. Moreover, since the t
In this study, we explore the transformer's ability to capture intra-relations among frames by augmenting the receptive field of models. Concretely, we propose a CycleGAN-based model with the transformer and investigate its ability in the emotional voice conversion task. In the training procedure, we adopt curriculum learning to gradually increase the frame length so that the model can see from the short segment till the entire speech. The proposed method was evaluated on the Japanese emotional
Open papers in the app to read, cite, and organize with AI.