[论文解读] LLM-based Smart Reply (LSR): Enhancing Collaborative Performance with ChatGPT-mediated Smart Reply System
本文提出一种基于LLM的智能回复(LSR)系统,利用ChatGPT在Slack中生成与上下文相关的回复,降低认知工作负荷并在模拟协作任务中提升工作表现与生产力。
Interactive user interfaces have increasingly explored AI's role in enhancing communication efficiency and productivity in collaborative tasks. The emergence of Large Language Models (LLMs) such as ChatGPT has revolutionized conversational agents, employing advanced deep learning techniques to generate context-aware, coherent, and personalized responses. Consequently, LLM-based AI assistants provide a more natural and efficient user experience across various scenarios. In this paper, we study how LLM models can be used to improve work efficiency in collaborative workplaces. Specifically, we present an LLM-based Smart Reply (LSR) system utilizing the ChatGPT to generate personalized responses in professional collaborative scenarios while adapting to context and communication style based on prior responses. Our two-step process involves generating a preliminary response type (e.g., Agree, Disagree) to provide a generalized direction for message generation, thus reducing response drafting time. We conducted an experiment where participants completed simulated work tasks involving a Dual N-back test and subtask scheduling through Google Calendar while interacting with co-workers. Our findings indicate that the proposed LSR reduces overall workload, as measured by the NASA TLX, and improves work performance and productivity in the N-back task. We also provide qualitative analysis based on participants' experiences, as well as design considerations to provide future directions for improving such implementations.
研究动机与目标
- 评估基于LLM的智能回复系统是否能在协作工作任务中降低认知工作负荷。
- 评估LSR对模拟工作场景中工作表现与生产力的影响。
- 探讨用户体验因素、信任、隐私以及AI辅助工作场所沟通的设计考量。
- 提供AI驱动协作工具的设计建议与未来方向。
提出的方法
- 三部分系统:将N-back任务作为模拟工作、Google Calendar用于子任务、并将Slack与LSR整合。
- 两步LSR工作流:生成初步回复类型(如同意/不同意)以指导信息生成,然后呈现三个AI生成的回复选项。
- 使用ChatGPT(GPT-3.5-turbo)基于最近十条消息生成回复;用户点击按钮发送AI生成的回复。
- 混合方法评估,包含定量指标(N-back准确率、每分钟消息数、NASA TLX)与定性数据(调查与半结构化访谈)。
- 参与者在在线仿真中与不同性格的同事(Jeff、Tony、Janine)互动,有/Lแจ вп LSR。
实验结果
研究问题
- RQ1RQ1: LSR是否影响协作工作场所中的工作表现、生产力和工作负荷?
- RQ2RQ2: 在协作工作中哪些关键因素影响用户体验?
主要发现
- LSR显著提升工作表现:平均N-back准确率从73.79%(SD 14.93)提升至79.37%(SD 8.96);差异为5.58%(p = 0.025)。
- LSR提高生产力:平均每分钟消息数提升40.36%(p = 3.74e-06)。
- NASA TLX结果显示LSR降低了心智负荷和时间负荷; Mental Demand(t = 2.7102,p = 0.0154)、Temporal Demand(t = 3.6794,p = 0.0020)改善,Performance呈边际变化(t = -2.5156,p = 0.0229)。
- 参与者报告工作流程更顺畅、恢复任务更快,且对AI生成回复的礼貌性/质量有正面感知,但也对上下文准确性与控制有所担忧。
- 设计考量包括控制与自动化之间的界面权衡、信任/隐私问题,以及为提高可用性需要可编辑的最终消息。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。