[Paper Review] AI Insights: A Case Study on Utilizing ChatGPT Intelligence for Research Paper Analysis
This study evaluates GPT-4 and GPT-3.5 for automating research paper analysis in scientific literature surveys, focusing on AI in breast cancer treatment. Using a corpus of 1,200+ papers from Google Scholar, PubMed, and Scopus, GPT-4 achieved 77.3% accuracy in category classification and 50% accuracy in scope detection, with expert-validated reasoning in 67% of cases.
This paper discusses the effectiveness of leveraging Chatbot: Generative Pre-trained Transformer (ChatGPT) versions 3.5 and 4 for analyzing research papers for effective writing of scientific literature surveys. The study selected the extit{Application of Artificial Intelligence in Breast Cancer Treatment} as the research topic. Research papers related to this topic were collected from three major publication databases Google Scholar, Pubmed, and Scopus. ChatGPT models were used to identify the category, scope, and relevant information from the research papers for automatic identification of relevant papers related to Breast Cancer Treatment (BCT), organization of papers according to scope, and identification of key information for survey paper writing. Evaluations performed using ground truth data annotated using subject experts reveal, that GPT-4 achieves 77.3\% accuracy in identifying the research paper categories and 50\% of the papers were correctly identified by GPT-4 for their scopes. Further, the results demonstrate that GPT-4 can generate reasons for its decisions with an average of 27\% new words, and 67\% of the reasons given by the model were completely agreeable to the subject experts.
Motivation & Objective
- To assess the effectiveness of GPT-3.5 and GPT-4 in automating research paper analysis for scientific literature surveys.
- To identify and classify research papers related to AI in breast cancer treatment (BCT) using LLMs.
- To evaluate the models' ability to detect the scope of BCT research papers compared to expert-annotated ground truth.
- To extract key information from research papers for survey paper writing using LLM-generated summaries.
- To identify limitations in LLM application for scholarly work, including data noise, response inconsistency, and API constraints.
Proposed method
- Constructed a taxonomy of BCT subdomains to guide paper collection and classification.
- Collected 1,200+ research papers from Google Scholar, PubMed, and Scopus, deduplicated to form a unified corpus.
- Used GPT-3.5 and GPT-4 to classify papers into BCT-related categories via prompts analyzing titles, abstracts, and content.
- Applied iterative prompt engineering to optimize classification, scope detection, and information extraction tasks.
- Evaluated model outputs against ground truth data annotated by subject-matter experts for category, scope, and reasoning quality.
- Performed information extraction on a sample paper using GPT-4 to retrieve background, methods, and key findings.
Experimental results
Research questions
- RQ1Can GPT-4 accurately classify research papers into relevant categories for AI in breast cancer treatment?
- RQ2How accurately can GPT-4 detect the scope of BCT research papers compared to expert-annotated ground truth?
- RQ3To what extent do GPT-4-generated reasoning statements align with expert judgments?
- RQ4What are the key limitations of using GPT models for academic literature analysis in real-world research workflows?
- RQ5How effective is GPT-4 in extracting structured information (e.g., objectives, methods, findings) from research papers?
Key findings
- GPT-4 achieved 77.3% accuracy in identifying the correct research paper category for BCT-related studies.
- GPT-4 correctly identified the scope of 50% of the research papers when compared to expert-annotated ground truth.
- Among scope detection results, 22% were intermediate matches, with 16% of cases showing broader scope and 4% narrower than expert-identified scope.
- GPT-4 generated reasoning for its decisions with an average of 27% new words compared to the input, indicating meaningful elaboration.
- 67% of the reasoning statements produced by GPT-4 were fully agreeable to subject-matter experts, indicating strong alignment with expert judgment.
- Key limitations included noisy data from Google Scholar and PubMed APIs, inconsistent LLM responses across iterations, and strict message limits in the GPT-4 API that hindered automation efficiency.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.