[Paper Review] Judgments of research co-created by generative AI: experimental evidence
This experimental study investigates public perceptions of research co-created with generative AI, finding that participants judge scientists who delegate tasks to large language models (LLMs) as less morally acceptable, less trustworthy, and more likely to produce lower-quality work than those delegating to human PhD students. The key result shows significant negative bias against AI-assisted research, with effect sizes ranging from d = -0.78 to d = -0.85 across moral acceptability, trust, and perceived quality.
The introduction of ChatGPT has fuelled a public debate on the use of generative AI (large language models; LLMs), including its use by researchers. In the current work, we test whether delegating parts of the research process to LLMs leads people to distrust and devalue researchers and scientific output. Participants (N=402) considered a researcher who delegates elements of the research process to a PhD student or LLM, and rated (1) moral acceptability, (2) trust in the scientist to oversee future projects, and (3) the accuracy and quality of the output. People judged delegating to an LLM as less acceptable than delegating to a human (d = -0.78). Delegation to an LLM also decreased trust to oversee future research projects (d = -0.80), and people thought the results would be less accurate and of lower quality (d = -0.85). We discuss how this devaluation might transfer into the underreporting of generative AI use.
Motivation & Objective
- To examine whether delegating research tasks to generative AI leads to negative judgments of researchers and their work.
- To compare public perceptions of AI-assisted research versus human-assisted research in terms of moral acceptability, trustworthiness, and perceived quality.
- To investigate whether such biases could lead to underreporting of AI use in academic research.
- To understand the psychological and social implications of integrating LLMs into scholarly workflows.
Proposed method
- Conducted an experimental study with 402 participants assessing judgments of researchers who delegated parts of a research process to either a PhD student or a large language model (LLM).
- Used a controlled scenario-based design where participants rated the moral acceptability, trustworthiness, and perceived accuracy/quality of research outputs.
- Employed standardized rating scales for moral acceptability, trust in oversight, and perceived accuracy and quality of results.
- Calculated Cohen’s d effect sizes to quantify differences in judgments between AI and human delegation conditions.
- Analyzed data using within-subjects and between-subjects comparisons to assess perception differences.
Experimental results
Research questions
- RQ1How do people judge the moral acceptability of researchers who delegate research tasks to generative AI versus human PhD students?
- RQ2To what extent does delegating to an LLM reduce perceived trust in a researcher’s ability to oversee future projects?
- RQ3How do perceptions of accuracy and quality differ when research is co-created with an LLM versus a human researcher?
- RQ4What are the potential consequences of these biases for the reporting and adoption of AI in academic research?
Key findings
- Participants judged delegating research tasks to an LLM as significantly less morally acceptable than delegating to a human PhD student, with a large effect size (d = -0.78).
- Trust in a researcher’s ability to oversee future projects was substantially lower when they used an LLM, with a large effect size (d = -0.80).
- Participants perceived the accuracy and quality of research outputs as lower when co-created with an LLM, with a large effect size (d = -0.85).
- The results suggest a strong societal bias against AI-assisted research, which may lead to underreporting of AI use in scholarly work.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.