[논문 리뷰] Investigating Affective Use and Emotional Well-being on ChatGPT
두 가지 병렬 연구—플랫폼상 3M 대화의 분석과 ~1,000명의 참가자를 대상으로 한 IRB 승인 RCT—은 ChatGPT의 정서적 사용이 정서적 행복에 어떻게 연결되는지 조사하며, 높은 사용량이 의존성과 음성의 미묘한 영향을 좌우함을 시사한다.
As AI chatbots see increased adoption and integration into everyday life, questions have been raised about the potential impact of human-like or anthropomorphic AI on users. In this work, we investigate the extent to which interactions with ChatGPT (with a focus on Advanced Voice Mode) may impact users' emotional well-being, behaviors and experiences through two parallel studies. To study the affective use of AI chatbots, we perform large-scale automated analysis of ChatGPT platform usage in a privacy-preserving manner, analyzing over 3 million conversations for affective cues and surveying over 4,000 users on their perceptions of ChatGPT. To investigate whether there is a relationship between model usage and emotional well-being, we conduct an Institutional Review Board (IRB)-approved randomized controlled trial (RCT) on close to 1,000 participants over 28 days, examining changes in their emotional well-being as they interact with ChatGPT under different experimental settings. In both on-platform data analysis and the RCT, we observe that very high usage correlates with increased self-reported indicators of dependence. From our RCT, we find that the impact of voice-based interactions on emotional well-being to be highly nuanced, and influenced by factors such as the user's initial emotional state and total usage duration. Overall, our analysis reveals that a small number of users are responsible for a disproportionate share of the most affective cues.
연구 동기 및 목표
- ChatGPT와의 상호작용이 네 가지 심리사회적 결과에 미치는 영향을 평가한다: 외로움, 사교성, 정서적 의존, 그리고 문제적 사용.
- 사용자 프라이버시를 보장하면서 자동 분류기를 사용해 대규모 플랫폼 내 대화를 분석하고 정서적 신호를 탐지한다.
- 모델 구성의 변화가 시간에 따른 사용자 행복에 어떤 영향을 미치는지 연구하기 위해 IRB 승인 무작위 대조시험을 수행한다.
- 정서적 신호를 주도하는 소수의 사용자 패턴과 음성 모달리티가 웰빙에 미치는 미묘한 영향을 식별한다.
제안 방법
- EmoClassifiersV1 (and EmoClassifiersV2) 개발하여 최상위 계층과 하위 분류기로 이중 구조의 정서적 신호를 탐지한다.
- Advanced Voice Mode 사용에 대한 파워유저 대 제어유저 코호트로 플랫폼 내 분석을 수행하고 설문조사(4,000명 이상 응답자) 를 실시한다.
- IRB-approved randomized controlled trial을 ~981 completers와 함께 모달리티(engaging/neutral voice vs text) 및 매일 작업을 변화시키는 nine conditions에서 28일에 걸쳐 수행한다.
- RCT에서의 31,857 conversations를 분석하여 사용자-모델 상호작용과 자기보고된 결과 간의 관계를 조사한다.
- 분류기를 정확한 상호작용별 레이블이 아닌 프라이버시를 보장하고 설문 응답과 상관된 기술적 도구로 간주한다.
실험 결과
연구 질문
- RQ1참여도 높은 음성 기반 챗봇 상호작용은 외로움, 사교성, 정서적 의존, 그리고 문제적 사용에 영향을 주는 데 있어 텍스트나 중립 음성보다 다른가?
- RQ2개인 대화 프롬프트가 ChatGPT를 사용할 때 비개인적이거나 열린 프롬프트보다 다른 웰빙 결과를 초래하는가?
- RQ3사용 지속시간과 초기 정서 상태가 ChatGPT의 웰빙 영향에 어떻게 조절하는가(RCT에서)?
주요 결과
- 매우 높은 사용량(상위 10%)은 자기보고된 정서적 의존 증가와 인지된 사교성 저하와 관련이 있다.
- 소수의 파워 유저가 대화의 정서 신호에 불균형적으로 기여한다.
- RCT에서 사용 지속시간을 통제하면 음성 모델 사용이 정서적 웰빙과 더 관련이 있는 경향이 있었으나, 더 긴 사용과 초기 외로움이 더 나쁜 결과를 예측했다.
- 자동 정서 분류기는 일반적으로 자기보고 설문 응답과 일치하며, 플랫폼 내 분석과 RCT 분석은 방법론적으로 상호 보완적이다.
- 대부분의 대화는 중립적이거나 작업 지향적이지만, 일부 사용자군은 채팅에서 빈번한 정서 신호를 보인다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.