Skip to main content
QUICK REVIEW

[論文レビュー] Harnessing the Power of LLMs: Evaluating Human-AI Text Co-Creation through the Lens of News Headline Generation

Zijian Ding, Alison Smith|arXiv (Cornell University)|Oct 16, 2023
Topic Modeling被引用数 4
ひとこと要約

本研究は、40名の参加者を対象に制御された実験により、ニュース見出し生成における人間とAIの共同作成を、誘導、選択、後処理の3つのインタラクションタイプで評価した。その結果、誘導と選択を組み合わせたアプローチが、最小限の作業量で最高品質の見出しを生み出すことが判明した。また、AIの支援がユーザーのコントロール感や信頼感を損なうことはなかった。

ABSTRACT

To explore how humans can best leverage LLMs for writing and how interacting with these models affects feelings of ownership and trust in the writing process, we compared common human-AI interaction types (e.g., guiding system, selecting from system outputs, post-editing outputs) in the context of LLM-assisted news headline generation. While LLMs alone can generate satisfactory news headlines, on average, human control is needed to fix undesirable model outputs. Of the interaction methods, guiding and selecting model output added the most benefit with the lowest cost (in time and effort). Further, AI assistance did not harm participants' perception of control compared to freeform editing.

研究の動機と目的

  • 異なる人間-AIインタラクション手法が、ニュース見出し生成における品質、作業負荷、および所有感に与える影響を調査すること。
  • AI支援がユーザーの作成プロセスにおけるコントロール感や信頼感を損なうかどうかを評価すること。
  • テキスト共同作成タスクにおいて、人間のコントロールとAIの効率性の最適なバランスを特定すること。
  • 実証的なユーザーインタラクションデータに基づき、LLMを搭載したライティングツールの設計指針を提示すること。
  • 今後のニュース生成分野におけるLLMの評価やRLHFファインチューニングを目的として、840件の人が評価した見出しのデータセットを公開すること。

提案手法

  • 編集経験を持つ40名の参加者を対象に、被験者間設計の制御実験を実施した。
  • 誘導、選択、後処理の3つのインタラクションタイプをサポートするプロトタイプのLLM駆動型見出し生成システムを用いた。
  • 手動作成、AI専用生成、および3つの支援付きバージョン(誘導、誘導+選択、誘導+選択+後処理)の合計5つの見出し生成条件を比較した。
  • 見出し品質の評価に、Taste-Attractiveness-Clarity-Truth(TACT)フレームワークを20名の専門家評価者を用いて採用した。
  • 心理的要因およびユーザビリティ要因を評価するため、コントロール感、信頼感、作業負荷に関する主観的フィードバックを収集した。
  • 定量的および定性的なデータを分析し、各条件間でのパフォーマンス、効率性、ユーザーエクスペリエンスの差を比較した。
Figure 1: Human-AI interactions for text generation can fall within a range of no human control effort (AI-only) to full human control effort (manual methods) Ding and Chan ( 2023 ) . Selecting from model outputs ( Selection ) alone provides less control (but also is easier) than when adding additio
Figure 1: Human-AI interactions for text generation can fall within a range of no human control effort (AI-only) to full human control effort (manual methods) Ding and Chan ( 2023 ) . Selecting from model outputs ( Selection ) alone provides less control (but also is easier) than when adding additio

実験結果

リサーチクエスチョン

  • RQ1TACT基準に基づくと、どの人間-AIインタラクション手法が最高品質のニュース見出しを生み出すか?
  • RQ2どの手法が、見出し品質を維持しつつ、ユーザーの認識される作業負荷と時間コストを最小限に抑えるか?
  • RQ3AI支援は、最終的な見出しに対するユーザーのコントロール感や信頼感にどのように影響するか?
  • RQ4誘導、選択、後処理の各インタラクションタイプが、見出し品質およびユーザーエクスペリエンスに与える影響はどのように比較されるか?
  • RQ5誘導と選択の組み合わせは、単独の手法や手動作成を上回るパフォーマンスを示すか?

主な発見

  • LLMが単独で生成した見出しは、人間が作成した見出しと平均的な品質が同等であったが、依然として誤りの修正が必要であった。
  • 誘導と選択の組み合わせが、最も高い品質の見出しを生み出し、時間的・作業的負荷が最小限であった。
  • 後処理および手動作成は、誘導+選択よりも著しく高い作業負荷を要したが、見出し品質に見合った恩恵は得られなかった。
  • 参加者たちは、AI専用生成や手動編集を含むすべての条件において、同程度のコントロール感と信頼感を報告した。
  • 誘導+選択が、品質、効率性、ユーザーエクスペリエンスのバランスを最適化する最適なインタラクションパターンとして浮き彫りになった。
  • 本研究では、今後のニュース生成分野におけるLLMの評価やRLHFファインチューニングを目的として、840件の人が評価した見出しのデータセットを公開した。
Figure 2: Interface for human-AI news headline co-creation for guidance + selection + post-editing condition: (A) news reading panel, (B) perspectives (keywords) selection panel (multiple keywords can be selected), (C) headline selection panel with post-editing capability, and (D) difficulty rating
Figure 2: Interface for human-AI news headline co-creation for guidance + selection + post-editing condition: (A) news reading panel, (B) perspectives (keywords) selection panel (multiple keywords can be selected), (C) headline selection panel with post-editing capability, and (D) difficulty rating

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。