[Paper Review] Human Evaluation of Interpretability: The Case of AI-Generated Music Knowledge
This paper proposes a human-centered evaluation framework for interpretability of AI-generated music rules, using student-written verbal interpretations of symbolic AI outputs. By collecting and grading free-text descriptions of histograms representing music theory rules, the study demonstrates that humans can successfully decode complex AI-discovered knowledge, revealing both the potential and challenges of interpretability in AI for the arts and humanities.
Interpretability of machine learning models has gained more and more attention among researchers in the artificial intelligence (AI) and human-computer interaction (HCI) communities. Most existing work focuses on decision making, whereas we consider knowledge discovery. In particular, we focus on evaluating AI-discovered knowledge/rules in the arts and humanities. From a specific scenario, we present an experimental procedure to collect and assess human-generated verbal interpretations of AI-generated music theory/rules rendered as sophisticated symbolic/numeric objects. Our goal is to reveal both the possibilities and the challenges in such a process of decoding expressive messages from AI sources. We treat this as a first step towards 1) better design of AI representations that are human interpretable and 2) a general methodology to evaluate interpretability of AI-discovered knowledge representations.
Motivation & Objective
- To develop a systematic method for evaluating interpretability of AI-generated knowledge in non-technical domains, particularly in music theory.
- To investigate whether humans can meaningfully interpret complex, symbolic AI outputs (e.g., probability histograms of musical features) through free-form verbal descriptions.
- To create a transparent, rubric-based grading process for qualitative human interpretations of AI rules, ensuring fairness and credibility.
- To assess the feasibility of using human-generated textual interpretations as a proxy for evaluating the interpretability of AI-discovered knowledge representations.
- To lay the foundation for a generalizable methodology applicable to interpretability evaluation in other domains beyond music.
Proposed method
- Designed a web-based AI system, MUS-ROVER, that learns music composition rules from sheet music using data-driven pattern discovery.
- Represented music rules as histograms of feature values (e.g., voice leading, chord inversions) derived from empirical probability distributions over musical structures.
- Administered a two-week homework assignment to CS+Music students, requiring them to write free-text interpretations of 25 AI-generated rules (11 one-gram, 14 two-gram rules).
- Provided a self-pretraining phase with detailed definitions of key terms (e.g., chord, window, basis feature, n-gram, histogram) to standardize understanding.
- Conducted a transparent discussion session with students and teaching staff to co-develop a revised, inclusive rubric based on consensus on key music-theoretic keywords.
- Manually graded student responses against the final rubric using a two-point-per-rule scoring system, focusing on qualitative insight rather than numerical or symbolic repetition.
Experimental results
Research questions
- RQ1Can humans successfully interpret AI-generated music rules represented as symbolic, numeric histograms through free-form textual descriptions?
- RQ2To what extent do human interpretations align with established music theory concepts when decoding AI-discovered rules?
- RQ3What are the main challenges in human interpretation of AI-generated knowledge when responses are open-ended and qualitative?
- RQ4How can a reliable, transparent, and fair rubric be constructed for evaluating qualitative human interpretations of AI outputs?
- RQ5Can this methodological framework be generalized to evaluate interpretability of AI knowledge in other domains beyond music?
Key findings
- The majority of students successfully interpreted the AI-generated music rules, demonstrating that symbolic AI outputs can be meaningfully decoded by humans with minimal domain-specific training.
- Students were able to identify known music theory concepts (e.g., voice leading, chord inversions) in the histograms, confirming that AI-discovered rules align with established musical principles.
- A significant portion of interpretations revealed novel or unexpected insights not directly tied to known rules, indicating that AI can uncover non-obvious patterns.
- The most common error was misinterpreting the direction of conditional probability in n-gram rules, highlighting a key challenge in understanding sequential dependencies in AI outputs.
- The collaborative rubric development process led to a more inclusive and credible evaluation framework, increasing transparency and fairness in grading qualitative responses.
- The study provides empirical evidence that human evaluation of free-text interpretations is a viable and informative method for assessing the interpretability of AI-generated knowledge representations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.