Skip to main content
QUICK REVIEW

[論文レビュー] Exploring ChatGPT's Ability to Rank Content: A Preliminary Study on Consistency with Human Preferences

Yunjie Ji, Yan Gong|arXiv (Cornell University)|Mar 14, 2023
Topic Modeling被引用数 14
ひとこと要約

本論文は、ChatGPT のゼロショット能力でモデル生成コンテンツを順位付けする能力を検証し、さまざまなプロンプトとモデルに渡る人間の好みと比較した。

ABSTRACT

As a natural language assistant, ChatGPT is capable of performing various tasks, including but not limited to article generation, code completion, and data analysis. Furthermore, ChatGPT has consistently demonstrated a remarkable level of accuracy and reliability in terms of content evaluation, exhibiting the capability of mimicking human preferences. To further explore ChatGPT's potential in this regard, a study is conducted to assess its ability to rank content. In order to do so, a test set consisting of prompts is created, covering a wide range of use cases, and five models are utilized to generate corresponding responses. ChatGPT is then instructed to rank the responses generated by these models. The results on the test set show that ChatGPT's ranking preferences are consistent with human to a certain extent. This preliminary experimental finding implies that ChatGPT's zero-shot ranking capability could be used to reduce annotation pressure in a number of ranking tasks.

研究の動機と目的

  • ChatGPT が複数のモデルからの回答を人間の好みに従って順位付けできるかを評価する。
  • ChatGPT ベースのランキングと人間の注釈との一貫性を、多様なユースケースで評価する。
  • ChatGPT のランキング能力が、ランキングタスクの注釈作業を削減できるかを検討する。

提案手法

  • NLPタスク、ブレインストーミング、生成、オープンQA、数学、コードを含むプロンプトのテストセットを構築する。
  • 各プロンプトにつき Bloomz 系列、Text-davinci-003、ChatGPT、InstrLLM を含む5モデルから回答を生成する。
  • 一問一答と一括プロンプトの2方式で、5つの回答をランキングするよう ChatGPT に依頼する。
  • 各プロンプトごとに129名の注釈者に回答をランク付けさせ、順位を平均化して人間の好みの注釈を収集する。
  • ChatGPTと人間のランキングのスピアマン相関とトップ1一貫性率で一致を測定する。
  • ユースケースとプロンプト間のモデル比較を分析し、長所と限界を特定する。

実験結果

リサーチクエスチョン

  • RQ1Can ChatGPT rank multiple model outputs in a way that is consistent with human preferences?
  • RQ2How does ranking performance vary between one-by-one and one-for-all prompting schemes?
  • RQ3Which use cases (NLP tasks, brainstorming, generation, open QA, math, code) show higher or lower alignment with human judgments?
  • RQ4How do different models (ChatGPT, Text-davinci-003, Bloomz variants, InstrLLM) compare in ranking quality from ChatGPT’s perspective?

主な発見

  • ChatGPT’s rankings show a non-trivial level of agreement with human preferences across prompts.
  • One-by-one prompting generally yields higher alignment with human preferences than one-for-all prompting.
  • Brainstorming and generation prompts show higher alignment; other use cases like NLP tasks and open QA show room for improvement.
  • ChatGPT outperforms several baselines (including Bloomz variants) in overall rankings, with some gaps in code and math prompts.
  • InstrLLM can beat some baselines on brainstorming, generation, and math, but lags in NLP tasks.
  • There is evidence that ChatGPT may be useful for bootstrap annotation in RLHF and for identifying use cases where performance is weak.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。