[Paper Review] JCoLA: Japanese Corpus of Linguistic Acceptability
This paper introduces JCoLA, a Japanese corpus of 10,020 linguistically annotated sentences with binary acceptability judgments, manually extracted from textbooks, handbooks, and journal articles. It evaluates 9 Japanese language models and finds that while some surpass human performance on in-domain (simpler) judgments, no model exceeds human performance on out-of-domain (theoretically complex) judgments, particularly struggling with long-distance syntactic dependencies like verbal agreement and NPI licensing.
Neural language models have exhibited outstanding performance in a range of downstream tasks. However, there is limited understanding regarding the extent to which these models internalize syntactic knowledge, so that various datasets have recently been constructed to facilitate syntactic evaluation of language models across languages. In this paper, we introduce JCoLA (Japanese Corpus of Linguistic Acceptability), which consists of 10,020 sentences annotated with binary acceptability judgments. Specifically, those sentences are manually extracted from linguistics textbooks, handbooks and journal articles, and split into in-domain data (86 %; relatively simple acceptability judgments extracted from textbooks and handbooks) and out-of-domain data (14 %; theoretically significant acceptability judgments extracted from journal articles), the latter of which is categorized by 12 linguistic phenomena. We then evaluate the syntactic knowledge of 9 different types of Japanese language models on JCoLA. The results demonstrated that several models could surpass human performance for the in-domain data, while no models were able to exceed human performance for the out-of-domain data. Error analyses by linguistic phenomena further revealed that although neural language models are adept at handling local syntactic dependencies like argument structure, their performance wanes when confronted with long-distance syntactic dependencies like verbal agreement and NPI licensing.
Motivation & Objective
- To address the lack of comprehensive syntactic evaluation datasets for Japanese language models.
- To create a linguistically rigorous corpus of acceptability judgments reflecting both textbook-level and theoretically significant syntactic phenomena.
- To evaluate the syntactic knowledge of 9 diverse Japanese language models on this new benchmark.
- To identify systematic weaknesses in neural models, especially regarding long-distance syntactic dependencies.
Proposed method
- The corpus was constructed by manually extracting 10,020 sentences from linguistics textbooks, handbooks, and journal articles.
- Sentences were split into in-domain (86%) and out-of-domain (14%) data, with the latter categorized into 12 linguistic phenomena.
- Binary acceptability judgments were assigned by expert linguists for each sentence.
- Nine different types of pre-trained Japanese language models were fine-tuned and evaluated on the JCoLA benchmark.
- Performance was measured using accuracy on acceptability prediction, with error analysis stratified by linguistic phenomenon.
- The study compared model performance against human judgments, particularly focusing on differences between in-domain and out-of-domain data.
Experimental results
Research questions
- RQ1Can Japanese language models achieve human-level or better performance on acceptability judgment tasks for simpler, textbook-level syntactic phenomena?
- RQ2Do Japanese language models generalize effectively to more complex, theoretically significant syntactic phenomena found in academic literature?
- RQ3Which linguistic phenomena pose the greatest challenge for current neural language models in Japanese?
- RQ4To what extent do language models capture long-distance syntactic dependencies such as verbal agreement and NPI licensing?
- RQ5How does the distribution of linguistic phenomena in the training data affect model performance?
Key findings
- Several language models surpassed human performance on the in-domain subset of JCoLA, which contains simpler acceptability judgments from textbooks and handbooks.
- No model exceeded human performance on the out-of-domain subset, which includes theoretically significant judgments from journal articles.
- Models performed well on local syntactic dependencies such as argument structure and filler-gap phenomena, which involve relatively short-range dependencies.
- Performance significantly declined on long-distance dependencies, particularly for verbal agreement and NPI/NCI licensing, which require tracking syntactic features across sentence spans.
- The high accuracy on binding was attributed to the high proportion of acceptable examples (93.1%) in that category, suggesting data imbalance affects performance.
- The skewed distribution of linguistic phenomena in JCoLA—especially the low number of sentences (around 10) for some phenomena—limits reliable evaluation and model generalization.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.