[Paper Review] Exploring ChatGPT's Ability to Rank Content: A Preliminary Study on Consistency with Human Preferences
The paper examines ChatGPT’s zero-shot ability to rank model-generated content and compares its rankings to human preferences across diverse prompts and models.
As a natural language assistant, ChatGPT is capable of performing various tasks, including but not limited to article generation, code completion, and data analysis. Furthermore, ChatGPT has consistently demonstrated a remarkable level of accuracy and reliability in terms of content evaluation, exhibiting the capability of mimicking human preferences. To further explore ChatGPT's potential in this regard, a study is conducted to assess its ability to rank content. In order to do so, a test set consisting of prompts is created, covering a wide range of use cases, and five models are utilized to generate corresponding responses. ChatGPT is then instructed to rank the responses generated by these models. The results on the test set show that ChatGPT's ranking preferences are consistent with human to a certain extent. This preliminary experimental finding implies that ChatGPT's zero-shot ranking capability could be used to reduce annotation pressure in a number of ranking tasks.
Motivation & Objective
- Assess whether ChatGPT can rank responses from multiple models according to human preferences.
- Evaluate consistency between ChatGPT-based rankings and human annotations across diverse use cases.
- Investigate whether ChatGPT’s ranking ability can reduce annotation effort in ranking tasks.
Proposed method
- Construct a test set of prompts spanning NLP tasks, brainstorming, generation, open QA, math, and code.
- Generate responses from five models per prompt including Bloomz variants, Text-davinci-003, ChatGPT, and InstrLLM.
- Ask ChatGPT to rank the five responses using one-by-one and one-for-all prompting prompts.
- Collect human preference annotations by having 129 annotators rank responses per prompt and averaging ranks.
- Measure agreement using Spearman correlation between ChatGPT and human rankings and top-1 consistency rate.
- Analyze model comparisons across use cases and prompts to identify strengths and limitations.
Experimental results
Research questions
- RQ1Can ChatGPT rank multiple model outputs in a way that is consistent with human preferences?
- RQ2How does ranking performance vary between one-by-one and one-for-all prompting schemes?
- RQ3Which use cases (NLP tasks, brainstorming, generation, open QA, math, code) show higher or lower alignment with human judgments?
- RQ4How do different models (ChatGPT, Text-davinci-003, Bloomz variants, InstrLLM) compare in ranking quality from ChatGPT’s perspective?
Key findings
- ChatGPT’s rankings show a non-trivial level of agreement with human preferences across prompts.
- One-by-one prompting generally yields higher alignment with human preferences than one-for-all prompting.
- Brainstorming and generation prompts show higher alignment; other use cases like NLP tasks and open QA show room for improvement.
- ChatGPT outperforms several baselines (including Bloomz variants) in overall rankings, with some gaps in code and math prompts.
- InstrLLM can beat some baselines on brainstorming, generation, and math, but lags in NLP tasks.
- There is evidence that ChatGPT may be useful for bootstrap annotation in RLHF and for identifying use cases where performance is weak.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.