[Paper Review] Does GPT-4 pass the Turing test?
The study evaluates GPT-4 in a public online Turing Test, finding GPT-4 prompts achieve up to 41% success vs. humans 63%, with significant prompt-driven variance and no clear link between interrogator demographics and accuracy.
We evaluated GPT-4 in a public online Turing test. The best-performing GPT-4 prompt passed in 49.7% of games, outperforming ELIZA (22%) and GPT-3.5 (20%), but falling short of the baseline set by human participants (66%). Participants' decisions were based mainly on linguistic style (35%) and socioemotional traits (27%), supporting the idea that intelligence, narrowly conceived, is not sufficient to pass the Turing test. Participant knowledge about LLMs and number of games played positively correlated with accuracy in detecting AI, suggesting learning and practice as possible strategies to mitigate deception. Despite known limitations as a test of intelligence, we argue that the Turing test continues to be relevant as an assessment of naturalistic communication and deception. AI models with the ability to masquerade as humans could have widespread societal consequences, and we analyse the effectiveness of different strategies and criteria for judging humanlikeness.
Motivation & Objective
- Assess whether GPT-4 can be mistaken for human in an online Turing Test.
- Compare GPT-4 against GPT-3.5 and ELIZA baselines across multiple prompts.
- Analyze how prompt design, strategies, and interrogator characteristics influence passing likelihood.
- Examine why the Turing Test remains relevant for studying naturalistic communication and deception.
Proposed method
- Implemented a two-player online Turing Test with an interrogator and a witness.
- Created 25 AI witnesses using GPT-3.5 or GPT-4 with varied prompts and parameters (temperature, persona, errors, and AI-vs-human framing).
- Implemented a chat interface with time/length limits, typing delays, and a decision log for verdicts and confidence.
- Conducted experiments on turingtest.live with random assignment to human or AI witnesses and an AI-explore/exploit model chooser for witnesses.
- Collected demographic data and analyzed examiner strategies, verdict reasons, and performance across prompts.
- Compared GPT-4 prompts against GPT-3.5 prompts and an ELIZA baseline.
Experimental results
Research questions
- RQ1What is GPT-4’s success rate in an online Turing Test relative to humans and baseline models (ELIZA, GPT-3.5)?
- RQ2How does prompt design affect GPT-4’s ability to pass as human in an interactive setting?
- RQ3Which interrogator strategies and participant demographics predict correct identification of AI vs human?
- RQ4What reasons do interrogators give for AI vs human verdicts, and what does this reveal about human-likeness cues?
- RQ5Does experience with the task (number of games) improve interrogator accuracy or reliability?
Key findings
- GPT-4’s best prompt achieved a 41% success rate, outperforming ELIZA (27%) and GPT-3.5 baselines (5–14%), but well below human performance (63%).
- There is substantial prompt-to-prompt variability in GPT-4 performance, with some prompts around 6–26% and others around 40–41%.
- Interrogator accuracy on AI witnesses was not predicted by demographics or prior interaction with LLMs; no learning effect within the experiment significantly affected accuracy.
- Linguistic style and socio-emotional cues were the top factors in AI-vs-human judgments, rather than evidence of genuine knowledge or reasoning.
- ELIZA, despite being simple, fooled human interrogators in 27% of games, highlighting the ELIZA effect and limits of the Turing Test as a measure of intelligence.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.