[论文解读] Investigating Affective Use and Emotional Well-being on ChatGPT
两项平行研究——基于平台上的3百万次对话分析以及一项获得IRB批准、约1000名参与者的随机对照试验(RCT)——探讨ChatGPT的情感性使用与情感幸福感之间的关系,发现高使用率与依赖性相关,并且语音模式对情感幸福感有微妙的影响。
As AI chatbots see increased adoption and integration into everyday life, questions have been raised about the potential impact of human-like or anthropomorphic AI on users. In this work, we investigate the extent to which interactions with ChatGPT (with a focus on Advanced Voice Mode) may impact users' emotional well-being, behaviors and experiences through two parallel studies. To study the affective use of AI chatbots, we perform large-scale automated analysis of ChatGPT platform usage in a privacy-preserving manner, analyzing over 3 million conversations for affective cues and surveying over 4,000 users on their perceptions of ChatGPT. To investigate whether there is a relationship between model usage and emotional well-being, we conduct an Institutional Review Board (IRB)-approved randomized controlled trial (RCT) on close to 1,000 participants over 28 days, examining changes in their emotional well-being as they interact with ChatGPT under different experimental settings. In both on-platform data analysis and the RCT, we observe that very high usage correlates with increased self-reported indicators of dependence. From our RCT, we find that the impact of voice-based interactions on emotional well-being to be highly nuanced, and influenced by factors such as the user's initial emotional state and total usage duration. Overall, our analysis reveals that a small number of users are responsible for a disproportionate share of the most affective cues.
研究动机与目标
- 评估与ChatGPT的互动如何影响四个社会心理结果:孤独感、社交/交往、情感依赖以及问题性使用。
- 在平台上对大规模对话进行分析,使用自动分类器检测情感线索,同时保护用户隐私。
- 开展一项获得IRB批准的随机对照试验,研究模型配置如何随时间影响用户福祉。
- 识别模式,表明只有一小部分用户会推动情感线索,且语音模态对福祉有细微影响。
提出的方法
- 开发 EmoClassifiersV1(及 EmoClassifiersV2)以检测情感线索,采用顶层与子分类器的两层结构。
- 对高使用者与对照用户队列在 Advanced Voice Mode 使用方面进行平台内分析,并开展用户调查(超过4,000名受访者)。
- 开展一项IRB批准的随机对照试验,约有981名完成者,涵盖九个条件,变量为模态(互动/引人入胜语音对比中性语音与文本)以及每日任务,持续28天。
- 分析来自RCT的31,857次对话,研究用户-模型互动与自我报告结果之间的关系。
- 将分类器视为描述性工具——保护隐私、与调查响应相关而非每次交互的精确标签。
实验结果
研究问题
- RQ1互动性强的语音聊天机器人交互是否在孤独感、社交/交往、情感依赖和问题性使用方面与文本或中性语音不同?
- RQ2在使用 ChatGPT 时,个性化对话提示是否比非个人或开放式提示导致不同的幸福感结果?
- RQ3使用时长和初始情绪状态如何调节 ChatGPT 对幸福感的影响(来自RCT)?
主要发现
- 非常高的使用量(前十百分位)与自我报告的情感依赖上升以及感知社交化程度下降相关。
- 少量高使用者对对话中的情感线索贡献不成比例。
- 在RCT中,控制使用时长后,语音模态的使用通常与更好的情感福祉相关,但更长的使用时间和更高的初始孤独感预测更差的结果。
- 自动化情感分类器通常与自我报告的调查响应一致,平台内分析与RCT分析在方法学上相互补充。
- 大多数对话是中性或任务导向的,但有一部分用户的聊天中频繁出现情感线索。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。