Skip to main content
QUICK REVIEW

[논문 리뷰] User Behavior Simulation with Large Language Model based Agents

Lei Wang, Jingsen Zhang|arXiv (Cornell University)|2023. 06. 05.
Topic Modeling인용 수 4
한 줄 요약

이 논문은 대규모 언어 모델(LLM) 기반의 에이전트 프레임워크를 제안하며, 추천 시스템과 소셜 네트워크와 같은 다양한 상호작용 환경에서 LLM의 세계 지식과 추론 능력을 활용하여 현실적인 사용자 행동을 시뮬레이션한다. 이 방법은 제로샷 설정에서도 인간 수준에 가까운 행동 유사도를 달성하여 정보 코니(정보 코니)와 순응성 같은 사회 현상을 고도로 현실감 있게 연구할 수 있도록 한다.

ABSTRACT

Simulating high quality user behavior data has always been a fundamental problem in human-centered applications, where the major difficulty originates from the intricate mechanism of human decision process. Recently, substantial evidences have suggested that by learning huge amounts of web knowledge, large language models (LLMs) can achieve human-like intelligence. We believe these models can provide significant opportunities to more believable user behavior simulation. To inspire such direction, we propose an LLM-based agent framework and design a sandbox environment to simulate real user behaviors. Based on extensive experiments, we find that the simulated behaviors of our method are very close to the ones of real humans. Concerning potential applications, we simulate and study two social phenomenons including (1) information cocoons and (2) user conformity behaviors. This research provides novel simulation paradigms for human-centered applications.

연구 동기 및 목표

  • 단순화된 의사결정 모델과 실제 데이터에 의존하는 전통적인 사용자 행동 시뮬레이션 방법의 한계를 해결하기 위해.
  • 실제 데이터 없이 LLM을 활용해 제로샷 사용자 행동 생성이 가능하도록 하여 시뮬레이션에서의 '닭과 계란' 문제를 해결하기 위해.
  • 다양한 맥락에서 사용자 행동 간 상호의존성을 포괄하는 다중 환경 시뮬레이션 프레임워크를 구축하기 위해.
  • 정보 코니와 순응성과 같은 잠재적인 사회 현상을 현실적인 LLM 기반 사용자 에이전트를 통해 연구하기 위해.
  • LLM 기반 에이전트를 활용해 인간 중심의 AI 응용 분야를 위한 새로운 시뮬레이션 패러다임을 수립하기 위해.

제안 방법

  • 추천 시스템과 소셜 미디어 플랫폼을 포함한 실제 사용자 상호작용 시나리오를 모방하는 샌드박스 환경을 설계하기 위해.
  • 자연어 추론과 세계 지식을 사용해 사용자 결정을 시뮬레이션하는 LLM 기반 에이전트를 훈련하고 배포하기 위해.
  • 핵심 행동에 대한 행동 템플릿을 구현: 영화 구매, 페이지 이동, 검색, 퇴장, 감정 생성, 대화 시작, 소셜 미디어 게시.
  • 프롬프트 엔지니어링을 활용해 에이전트가 맥락에 적절한 방식으로 행동하도록 유도하여 관찰된 사용자 행동 패tern과의 일관성을 확보하기 위해.
  • 다단계 상호작용 로직을 통합해 환경 간 동적인, 변화하는 사용자 상호작용을 시뮬레이션하기 위해.
  • 미리 훈련된 LLM을 활용해 실제 사용자 데이터에 대한 피팅 없이도 제로샷 시뮬레이션을 가능하게 하기 위해.
Figure 1 : a , A brief running process of the simulator. b , The agent framework, which includes a profile module, a memory module, and an action module. c , Key characters of the simulator. Different agents behave in a round-by-round manner based on Pareto distribution, where, in each round, only a
Figure 1 : a , A brief running process of the simulator. b , The agent framework, which includes a profile module, a memory module, and an action module. c , Key characters of the simulator. Different agents behave in a round-by-round manner based on Pareto distribution, where, in each round, only a

실험 결과

연구 질문

  • RQ1LLM 기반 에이전트는 다중 환경 설정에서 실제 인간의 행동과 구분이 가지 않는 사용자 행동을 생성할 수 있는가?
  • RQ2실제 세계 데이터에 접근하지 못한 상태에서 LLM 에이전트는 정보 코니와 순응성과 같은 복잡한 사회 현상을 어느 정도 정확하게 시뮬레이션할 수 있는가?
  • RQ3실제 데이터에 의존하는 전통적 시뮬레이션 기법과 비교해 LLM 기반 에이전트는 제로샷 시뮬레이션에서 어떤 성능을 보이는가?
  • RQ4LLM 기반 에이전트는 예를 들어 시청, 대화, 게시와 같은 다양한 상호작용 모odalities 간에 행동 일관성과 맥락 적절성을 유지할 수 있는가?
  • RQ5공동 시뮬레이션 환경에 다수의 LLM 에이전트를 배치했을 때 어떤 잠재적인 사회 역학적 동작이 관찰될 수 있는가?

주요 결과

  • LLM 기반 에이전트의 시뮬레이션된 행동은 추천 시스템과 소셜 네트워크 양쪽 모두에서 실제 인간의 행동과 매우 유사하며 현실감이 높다.
  • 이 프레임워크는 제로샷 시뮬레이션을 가능하게 하여 실제 세계의 훈련 데이터가 전혀 필요 없이 유의미한 사용자 행동을 생성할 수 있다.
  • 에이전트는 정보 코니와 사용자 순응성과 같은 복잡한 사회 현상을 성공적으로 시뮬레이션하여 현실적인 사회 역학의 탄생을 보여주었다.
  • 에이전트는 사회적 신호에 대한 적절한 반응, 감정 반영, 타겟된 콘텐츠 검색 등 맥락 인식 기반의 의사결정을 보였다.
  • 이 시뮬레이션 프레임워크는 다중 모odal 및 다중 환경 상호작용을 지원하여 플랫폼 간 사용자 행동을 현실적으로 모델링할 수 있었다.
  • 이 방법은 단순 모델과 실제 데이터 의존성에 기반한 전통적 시뮬레이션 기법보다 뛰어난 성능을 보이며 인간 중심의 AI 연구 분야에 더 넓은 적용 가능성을 제공한다.
Figure 2 : Evaluation on the believability of the simulated user behaviors. a , Evaluation on the recommendation behaviors based on different $(a,b)$ ’s (discrimination capability). b , Evaluation on the recommendation behaviors based on different $N$ ’s (generation capability). c , Evaluation on th
Figure 2 : Evaluation on the believability of the simulated user behaviors. a , Evaluation on the recommendation behaviors based on different $(a,b)$ ’s (discrimination capability). b , Evaluation on the recommendation behaviors based on different $N$ ’s (generation capability). c , Evaluation on th

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.