[论文解读] Thinking Assistants: LLM-Based Conversational Assistants that Help Users Think By Asking rather than Answering
本文介紹了 *Thinking Assistants*,這是一種基於大語言模型(LLM)的對話代理,透過提問而非直接提供答案來促進使用者的深度反思。該系統,* Sys*,透過個人化、類似導師的對話,協助潛在的畢業學生明確研究興趣,結果顯示使用者滿意度達65%,且在討論個人研究時的互動時間是查詢教授資訊時的兩倍。
Many AI systems focus solely on providing solutions or explaining outcomes. However, complex tasks like research and strategic thinking often benefit from a more comprehensive approach to augmenting the thinking process rather than passively getting information. We introduce the concept of "Thinking Assistant", a new genre of assistants that help users improve decision-making with a combination of asking reflection questions based on expert knowledge. Through our lab study (N=80), these Large Language Model (LLM) based Thinking Assistants were better able to guide users to make important decisions, compared with conversational agents that only asked questions, provided advice, or neither. Based on the results, we develop a Thinking Assistant in academic career development, determining research trajectory or developing one's unique research identity, which requires deliberation, reflection and experts' advice accordingly. In a longitudinal deployment with 223 conversations, participants responded positively to approximately 65% of the responses. Our work proposes directions for developing more effective LLM agents. Rather than adhering to the prevailing authoritative approach of generating definitive answers, LLM agents aimed at assisting with cognitive enhancement should prioritize fostering reflection. They should initially provide responses designed to prompt thoughtful consideration through inquiring, followed by offering advice only after gaining a deeper understanding of the user's context and needs.
研究动机与目标
- 解決高風險決策(如研究生院申請)中自我探索的挑戰,因導師指導往往難以取得或令人感到威嚇。
- 透過從資訊傳遞轉向反思對話,減少使用者在決策過程中的挫折感,特別是針對不確定或身份形成中的選擇。
- 設計一種對話助理,透過主動參與與智慧反饋,促進使用者對研究興趣的承諾,模擬導師角色。
- 評估思考助理模型是否在學術決策情境中,相較於傳統基於問答的聊天機器人,能提升使用者參與度與滿意度。
- 探討反思對話與使用者對明確答案(如「我有機會嗎?」)期望之間的張力,特別是在申請成功機率等議題上。
提出的方法
- 該系統 *\Sys* 使用微調過的 GPT-4,訓練資料來源為參與的 HCI 教授所提供的精選資料,包括研究演變歷程、指導風格與核心特質。
- 系統運作於兩種模式:『探問模式』,透過提問深化使用者思考;『回覆模式』,提供關於教授的實質資訊。
- 助理採用動態回覆策略,多數回合均以追問作結,以維持對話流暢並促進反思。
- 另設一項次級安全機器人,用於驗證事實性陳述,特別是關於論文與課程細節,以防止錯誤資訊傳播。
- 系統透過與參與教授的反覆設計與驗證,確保其專業知識與指導方式得到準確呈現。
- 使用者對話紀錄被記錄並分析,共計 223 次互動,用以評估參與度、滿意度與回覆模式。

实验结果
研究问题
- RQ1強調提問而非回答的思考助理,如何影響在研究生院申請情境中使用者的參與度與滿意度?
- RQ2使用者在做出高風險決策(如研究生院申請)時,對反思對話與直接答案的偏好程度如何?
- RQ3個人研究興趣的披露在多大程度上影響與思考助理互動的品質與深度?
- RQ4當使用者期望獲得明確答案(如「我有機會嗎?」)時,思考助理的限制為何?
- RQ5思考助理如何在保持反思性支架的同時,有效整合事實資訊檢索,而不損害使用者信任?
主要发现
- 使用者對回覆的滿意度達 65%,且在討論自身研究興趣時,參與度顯著高於僅查詢教授資訊的情境。
- 涉及個人研究披露的對話,平均每回合達六則訊息,顯示更深層參與;而僅獲取資訊的查詢則較短,且滿意度較低。
- 期望獲得明確答案(如「我有機會嗎?」)的使用者,認為反思模式令人不滿,突顯使用者期望與助理設計之間的落差。
- 即使僅是基本的事實查詢(如論文推薦),使用者仍認為有幫助,儘管公開資訊已存在,顯示經篩選、易取得的回覆具有價值。
- 系統的安全機器人成功修正了事實性錯誤,顯示在基於大語言模型的助理中,驗證層級至關重要。
- 雙模式設計(探問 vs. 回覆)提供了彈性,但使用者在助理優先採用反思模式而非提供資訊時,難以適應。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。