[Paper Review] ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor
The paper audits GPT-3.5 and GPT-4 across four scientific roles—librarian, ethicist, data generator, and data predictor—finding improving performance in some areas (e.g., hallucination reduction, ethics detection) but limited ability to predict novel data.
How good a research scientist is ChatGPT? We systematically probed the capabilities of GPT-3.5 and GPT-4 across four central components of the scientific process: as a Research Librarian, Research Ethicist, Data Generator, and Novel Data Predictor, using psychological science as a testing field. In Study 1 (Research Librarian), unlike human researchers, GPT-3.5 and GPT-4 hallucinated, authoritatively generating fictional references 36.0% and 5.4% of the time, respectively, although GPT-4 exhibited an evolving capacity to acknowledge its fictions. In Study 2 (Research Ethicist), GPT-4 (though not GPT-3.5) proved capable of detecting violations like p-hacking in fictional research protocols, correcting 88.6% of blatantly presented issues, and 72.6% of subtly presented issues. In Study 3 (Data Generator), both models consistently replicated patterns of cultural bias previously discovered in large language corpora, indicating that ChatGPT can simulate known results, an antecedent to usefulness for both data generation and skills like hypothesis generation. Contrastingly, in Study 4 (Novel Data Predictor), neither model was successful at predicting new results absent in their training data, and neither appeared to leverage substantially new information when predicting more versus less novel outcomes. Together, these results suggest that GPT is a flawed but rapidly improving librarian, a decent research ethicist already, capable of data generation in simple domains with known characteristics but poor at predicting novel patterns of empirical data to aid future experimentation.
Motivation & Objective
- Evaluate GPT-3.5 and GPT-4 as a Research Librarian by testing bibliography quality and hallucination rates.
- Evaluate GPT-3.5 and GPT-4 as a Research Ethicist by measuring detection and correction of flawed research practices.
- Assess GPT-3.5 and GPT-4 as Data Generators by examining bias replication and ability to simulate known results.
- Assess GPT-3.5 and GPT-4 as Novel Data Predictors by testing predictions on unseen, real-world data patterns.
Proposed method
- Study 1 (Librarian): generate 1,000 references (20 per topic across 25 psychology topics) and rate for correctness, completeness, relevance, and citation count.
- Study 2 (Ethicist): present 18 vignettes (blatant and subtle) of flawed protocols and rate GPT responses on ethical/reflexive quality across 216 interactions.
- Study 3 (Data Generator): have GPT estimate word-embedding-like associations and replicate known bias patterns in a four-domain WEAT-inspired assessment.
- Study 4 (Novel Data Predictor): task GPT with predicting country-level attitudes ( implicit vs explicit ) using Project Implicit data to assess novelty and predictive validity.
- Quantitative analyses include logistic regressions, Cronbach alpha reliability, and correlation analyses against real-world data.
Experimental results
Research questions
- RQ1Can GPT reliably compile a comprehensive and accurate bibliography without hallucinating?
- RQ2How well does GPT detect and address ethical issues and p-hacking-like practices in research protocols?
- RQ3To what extent can GPT simulate known data patterns (biases, stereotypes) and generate plausible data?
- RQ4Is GPT capable of predicting novel empirical patterns outside its training data, and how does performance differ between GPT-3.5 and GPT-4?
- RQ5What are the limitations and trajectory of GPT's utility as a general scientific assistant?
Key findings
- GPT-3.5 hallucinated 36.0% of references; GPT-4 hallucinated 5.4% of references, with GPT-4 showing improved candor about fiction (acknowledging fiction) 84.3% of the time for fictional references vs 12.2% for GPT-3.5.
- GPT-4 outperformed GPT-3.5 on ethical vignette responses, averaging 8.86/10 (blatant) and 7.26/10 (subtle) versus 5.39/10 and 4.05/10 respectively.
- GPT reliably replicates known bias patterns in data generation (e.g., WEAT-like results) and can simulate established outcomes, supporting its use for pilot data generation and hypothesis generation.
- In the novel data predictor task, GPT-3.5 and GPT-4 showed limited ability to predict highly novel data; correlations with real-world results varied and declined for novel implicit attitudes.
- Prompting with data-ethics themes boosted response quality, with ethical-primed prompts yielding higher-quality outputs than non-ethics prompts.
- Overall, GPT is a flawed but improving librarian, a reasonable ethics tool, capable of simple-domain data generation, but poor at predicting novel empirical patterns.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.