[Paper Review] Application of NotebookLM, a Large Language Model with Retrieval-Augmented Generation, for Lung Cancer Staging
This study evaluates NotebookLM, a retrieval-augmented generation (RAG) large language model, for lung cancer staging using Japan's 8th edition lung cancer staging guidelines as reliable external knowledge (REK). NotebookLM achieved 86% diagnostic accuracy in staging 100 fictional cases—significantly outperforming GPT-4o (39% with REK) and demonstrating 95% reference search accuracy, enabling radiologists to verify responses and reduce hallucinations.
Purpose: In radiology, large language models (LLMs), including ChatGPT, have recently gained attention, and their utility is being rapidly evaluated. However, concerns have emerged regarding their reliability in clinical applications due to limitations such as hallucinations and insufficient referencing. To address these issues, we focus on the latest technology, retrieval-augmented generation (RAG), which enables LLMs to reference reliable external knowledge (REK). Specifically, this study examines the utility and reliability of a recently released RAG-equipped LLM (RAG-LLM), NotebookLM, for staging lung cancer. Materials and methods: We summarized the current lung cancer staging guideline in Japan and provided this as REK to NotebookLM. We then tasked NotebookLM with staging 100 fictional lung cancer cases based on CT findings and evaluated its accuracy. For comparison, we performed the same task using a gold-standard LLM, GPT-4 Omni (GPT-4o), both with and without the REK. Results: NotebookLM achieved 86% diagnostic accuracy in the lung cancer staging experiment, outperforming GPT-4o, which recorded 39% accuracy with the REK and 25% without it. Moreover, NotebookLM demonstrated 95% accuracy in searching reference locations within the REK. Conclusion: NotebookLM successfully performed lung cancer staging by utilizing the REK, demonstrating superior performance compared to GPT-4o. Additionally, it provided highly accurate reference locations within the REK, allowing radiologists to efficiently evaluate the reliability of NotebookLM's responses and detect possible hallucinations. Overall, this study highlights the potential of NotebookLM, a RAG-LLM, in image diagnosis.
Motivation & Objective
- To assess the diagnostic accuracy and reliability of NotebookLM, a retrieval-augmented generation (RAG) large language model, in lung cancer staging.
- To evaluate whether providing a standardized, up-to-date clinical guideline as reliable external knowledge (REK) improves LLM performance in medical image diagnosis.
- To compare NotebookLM’s performance against GPT-4 Omni (GPT-4o), both with and without REK, in a controlled staging task.
- To examine the traceability of responses by analyzing how accurately NotebookLM cites locations within the provided REK.
- To explore the potential of RAG-LLMs as reliable diagnostic support tools in radiology, particularly in reducing hallucinations through source attribution.
Proposed method
- Constructed 100 fictional lung cancer cases with CT findings and corresponding TNM classifications based on Japan’s 8th edition lung cancer staging guideline.
- Summarized the official Japanese lung cancer staging guideline into a structured, machine-readable reliable external knowledge (REK) document.
- Used NotebookLM with the REK to generate TNM classifications for each case, leveraging its retrieval-augmented generation (RAG) mechanism to ground responses in the REK.
- Evaluated GPT-4o under two conditions: with and without the same REK, using the same prompt and case set for direct comparison.
- Measured diagnostic accuracy for T, N, M factors and overall TNM classification, and assessed reference search accuracy by verifying if NotebookLM cited correct REK passages.
- Validated results through consensus review by multiple radiologists, with final confirmation by a senior radiologist.
Experimental results
Research questions
- RQ1Can NotebookLM, a RAG-equipped LLM, achieve higher diagnostic accuracy than traditional LLMs like GPT-4o in lung cancer staging when provided with a standardized clinical guideline as REK?
- RQ2How accurately can NotebookLM locate and cite relevant sections within the provided REK, and does this enhance the traceability and trustworthiness of its responses?
- RQ3Does the inclusion of REK significantly improve the diagnostic performance of LLMs in a clinical staging task, and if so, to what extent?
- RQ4What types of errors persist in RAG-LLMs like NotebookLM despite source grounding, and how do they compare to errors in non-RAG models?
- RQ5To what extent can RAG-LLMs like NotebookLM be trusted as diagnostic support tools in radiology, given their ability to cite sources and reduce hallucinations?
Key findings
- NotebookLM achieved 86% diagnostic accuracy in staging 100 fictional lung cancer cases using the provided REK, significantly outperforming GPT-4o, which scored 39% with REK and 25% without.
- NotebookLM demonstrated 95% search accuracy in identifying and citing correct locations within the REK, enabling radiologists to verify the basis of its responses.
- The diagnostic accuracy of GPT-4o improved slightly with REK (39%) compared to without (25%), but remained substantially lower than NotebookLM’s performance.
- NotebookLM’s superior performance is attributed to its RAG architecture, which restricts generation to the provided REK and ensures source grounding.
- Despite high accuracy and traceability, NotebookLM still made errors, primarily due to incorrect numerical comparisons—similar to issues observed in GPT-4o.
- The study highlights that RAG-LLMs like NotebookLM can reduce hallucinations and improve reliability in clinical LLM applications, especially when combined with authoritative REK.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.