QUICK REVIEW
[Paper Review] Causal-Discovery Performance of ChatGPT in the context of Neuropathic Pain Diagnosis
Ruibo Tu, Chao Ma|arXiv (Cornell University)|Jan 24, 2023
TL;DR
The paper evaluates how well ChatGPT answers causal-discovery questions in neuropathic pain diagnosis using a sampled benchmark, finding limited and inconsistent performance with notable false negatives and language/concept limitations.
ABSTRACT
ChatGPT has demonstrated exceptional proficiency in natural language conversation, e.g., it can answer a wide range of questions while no previous large language models can. Thus, we would like to push its limit and explore its ability to answer causal discovery questions by using a medical benchmark (Tu et al. 2019) in causal discovery.
Motivation & Objective
- Assess ChatGPT's ability to answer causal-discovery questions in a medical domain (neuropathic pain).
- Analyze how ChatGPT answers questions about causal relationships using meta-information rather than observational data.
- Identify limitations, consistency issues, and language/concept understanding gaps in ChatGPT's responses.
Proposed method
- Create a benchmark by sampling 50 true causal pairs and 50 false pairs from a neuropathic pain causal map.
- Format questions as 'X causes Y. Answer true or false' using variable pairs.
- Query ChatGPT on whether the stated relationship is true or false.
- Analyze responses for precision, recall, and F-score.
- Inspect qualitative and quantitative errors to identify patterns of limitations.
Experimental results
Research questions
- RQ1Can ChatGPT reliably identify true causal relationships in a neuropathic pain causal map from meta-information alone?
- RQ2What are the common error modes (e.g., false negatives, language/region misinterpretation) in ChatGPT's causal judgments?
- RQ3How consistent are ChatGPT's answers across repeated queries and across languages or terminologies?
Key findings
- ChatGPT shows high precision but low recall in the tested causal judgments.
- There are no false positives in the reported confusion matrix, but several false negatives reduce overall performance.
- Answers vary across days, indicating instability and potential model-update effects on responses.
- ChatGPT struggles with regional anatomy descriptions (e.g., T12 involvement) and with non-English terms, contributing to misclassifications.
- Some queries yield reasonable or non-trivial correct responses, suggesting potential complementary use for causal discovery tasks.
- Overall, current ChatGPT performance in causal discovery is limited and not reliable for definitive causal claims.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.