[Paper Review] Can ChatGPT Really Understand Modern Chinese Poetry?
The paper presents ECUMP, a framework to evaluate ChatGPT’s understanding of modern Chinese poetry, showing 73% alignment with original poets’ intents across 48 poems, with weaker performance on poeticity.
ChatGPT has demonstrated remarkable capabilities on both poetry generation and translation, yet its ability to truly understand poetry remains unexplored. Previous poetry-related work merely analyzed experimental outcomes without addressing fundamental issues of comprehension. This paper introduces a comprehensive framework for evaluating ChatGPT's understanding of modern poetry. We collaborated with professional poets to evaluate ChatGPT's interpretation of modern Chinese poems by different poets along multiple dimensions. Evaluation results show that ChatGPT's interpretations align with the original poets' intents in over 73% of the cases. However, its understanding in certain dimensions, particularly in capturing poeticity, proved to be less satisfactory. These findings highlight the effectiveness and necessity of our proposed framework. This study not only evaluates ChatGPT's ability to understand modern poetry but also establishes a solid foundation for future research on LLMs and their application to poetry-related tasks.
Motivation & Objective
- Identify five dimensions essential for understanding modern poetry (content, expression methods, thought & emotion, modernity, poeticity) with expert input.
- Develop a prompt design to elicit multi-dimensional poem interpretations from ChatGPT.
- Compare ChatGPT interpretations with evaluations by professional poets to establish a ground truth.
- Provide an evaluation framework and evidence to guide future LLM-based poetry tasks and research.
Proposed method
- Define five poetry-understanding dimensions grounded in poetry theory and expert input.
- Design and optimize a ChatGPT prompt to interpret modern poetry across those dimensions (Content, Expression Methods, Thought & Emotion, Modernity, Poeticity).
- Assemble a 48-poem dataset (Com-Poetry and Spe-Poetry) from six professional poets for interpretation tasks.
- Use GPT-4 (gpt-4-0125) with fixed generation settings to produce interpretations across dimensions.
- Obtain evaluation from original poets using a 0–100 scale for four dimensions and 0/50/100 for Poeticity, plus a parallel LLM-based evaluation.

Experimental results
Research questions
- RQ1Does ChatGPT truly understand modern Chinese poetry across predefined dimensions?
- RQ2How well do ChatGPT interpretations align with the original poets’ intents across different poem types (Com-Poetry vs Spe-Poetry)?
- RQ3Which dimensions are most challenging for ChatGPT to capture (e.g., poeticity vs imagery)?
Key findings
- GPT-4’s interpretations align with the original poets’ intents in over 73% of cases across dimensions.
- Imagery understanding is strongest for Com-Poetry, with an average score of 81.18.
- For Spe-Poetry, strengths include rhetorical techniques (88.75), rhythm (82.50), and modernity (82.50).
- Poeticity is the weakest dimension for GPT-4, with poor identification of the most poetic sentence (Table shows many 0/50/100 outcomes).
- Human poets’ evaluations show higher reliability than automatic LLM evaluations for poetry understanding.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.