Skip to main content
QUICK REVIEW

[论文解读] WEMAC: Women and Emotion Multi-modal Affective Computing dataset

José A. Miranda, Esther Rituerto-González|arXiv (Cornell University)|Mar 1, 2022
Emotion and Mood Recognition被引用 5
一句话总结

本论文介绍了WEMAC,这是一个新颖的多模态数据集,包含47名女性在实验室环境中通过虚拟现实刺激诱发情绪反应后获得的生理、语音和自报告情绪数据。该数据集支持针对女性群体的情感计算研究,特别关注恐惧及其他与基于性别暴力预防相关的情绪,通过音频-视觉特征的后期融合,在二分类恐惧/非恐惧任务中实现了67.59的F1分数。

ABSTRACT

Among the seventeen Sustainable Development Goals (SDGs) proposed within the 2030 Agenda and adopted by all the United Nations member states, the Fifth SDG is a call for action to turn Gender Equality into a fundamental human right and an essential foundation for a better world. It includes the eradication of all types of violence against women. Within this context, the UC3M4Safety research team aims to develop Bindi. This is a cyber-physical system which includes embedded Artificial Intelligence algorithms, for user real-time monitoring towards the detection of affective states, with the ultimate goal of achieving the early detection of risk situations for women. On this basis, we make use of wearable affective computing including smart sensors, data encryption for secure and accurate collection of presumed crime evidence, as well as the remote connection to protecting agents. Towards the development of such system, the recordings of different laboratory and into-the-wild datasets are in process. These are contained within the UC3M4Safety Database. Thus, this paper presents and details the first release of WEMAC, a novel multi-modal dataset, which comprises a laboratory-based experiment for 47 women volunteers that were exposed to validated audio-visual stimuli to induce real emotions by using a virtual reality headset while physiological, speech signals and self-reports were acquired and collected. We believe this dataset will serve and assist research on multi-modal affective computing using physiological and speech information.

研究动机与目标

  • 为解决情感计算中缺乏性别平衡且具现实感的情绪数据集,特别是针对女性和与恐惧相关的情绪状态的问题。
  • 支持开发用于早期检测基于性别暴力(GBV)风险情境的智能可穿戴系统。
  • 提供一个公开可用、经过伦理审查的数据集,包含生理、语音和自报告数据,用于多模态情绪分析。
  • 支持在基于性别暴力预防背景下,利用多模态融合、迁移学习和注意力机制进行情绪识别的研究。
  • 通过聚焦女性情绪反应和生理激活模式,推动情感计算中的性别平等。

提出的方法

  • 参与者通过虚拟现实头戴设备接收经过验证的音视频刺激,在受控的实验室环境中诱发真实情绪。
  • 使用可穿戴传感器和标准硬件采集生理信号(ECG、EDA、呼吸、PPG)和语音信号。
  • 语音信号通过语音活动检测(VAD)处理,并利用librosa、openSMILE和DeepSpectrum工具包提取低级和高级特征。
  • 共提取了12,882个特征和嵌入,包括38个基于MFCC的特征、88个eGeMAPS特征、6,373个ComParE特征以及6,144维的DeepSpectrum嵌入。
  • 从在AudioSet上预训练的VGG-19网络中提取了128维的VGGish嵌入。
  • 为保护隐私,所有原始语音数据均被排除;仅发布处理后的特征和嵌入,并附带开源信号处理代码。

实验结果

研究问题

  • RQ1如何有效捕捉和处理多模态生理和语音信号,以表征女性的真实情绪状态?
  • RQ2在性别特定背景下,音频-视觉融合在检测恐惧与非恐惧状态方面能提升多少性能?
  • RQ3WEMAC数据集能否支持开发稳健且实时的情绪识别系统,用于基于性别暴力的风险检测?
  • RQ4在受控的实验室环境中,女性对刺激的性别特异性情绪反应如何比较?
  • RQ5基线模型在这一新颖的、面向女性的情感计算数据集上的表现如何?

主要发现

  • WEMAC数据集包含47名女性在受控实验室环境中,通过虚拟现实诱发刺激获得的生理、语音和自报告情绪数据。
  • 从语音信号中提取了共计12,882个特征和嵌入,包括38个MFCC基特征、88个eGeMAPS特征、6,373个ComParE特征以及6,144维的DeepSpectrum特征。
  • 官方基线模型在使用后期融合的音频-视觉刺激层级上,实现了67.59的F1分数,用于二分类恐惧/非恐惧任务。
  • 该数据集已公开发布,包含处理后的特征和开源信号处理代码,以确保隐私保护与可复现性。
  • 该数据集是更大规模研究的一部分,计划纳入额外的57名非GBV女性,并持续收集GBV幸存者的数据,以提升生态效度。
  • 研究团队计划通过Bindi——一种用于实时GBV风险检测的网络物理人工智能系统——将该系统扩展至真实世界可穿戴设备部署。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。