[论文解读] Grammatical cues to subjecthood are redundant in a majority of simple clauses across languages
本研究调查了30种语言中语法线索(如词序、词形变化)在识别主语时的冗余性。通过在英语和俄语中进行行为实验,并利用神经网络分类器对跨语言语料库数据进行分析,发现大多数简单及及物小句中,语法线索具有冗余性——即使没有这些线索,主语的识别准确率仍可达87%–89%,表明意义和合情合理性通常足以支持主语识别。
Grammatical cues are sometimes redundant with word meanings in natural language. For instance, English word order rules constrain the word order of a sentence like "The dog chewed the bone" even though the status of "dog" as subject and "bone" as object can be inferred from world knowledge and plausibility. Quantifying how often this redundancy occurs, and how the level of redundancy varies across typologically diverse languages, can shed light on the function and evolution of grammar. To that end, we performed a behavioral experiment in English and Russian and a cross-linguistic computational analysis measuring the redundancy of grammatical cues in transitive clauses extracted from corpus text. English and Russian speakers (n=484) were presented with subjects, verbs, and objects (in random order and with morphological markings removed) extracted from naturally occurring sentences and were asked to identify which noun is the subject of the action. Accuracy was high in both languages (~89% in English, ~87% in Russian). Next, we trained a neural network machine classifier on a similar task: predicting which nominal in a subject-verb-object triad is the subject. Across 30 languages from eight language families, performance was consistently high: a median accuracy of 87%, comparable to the accuracy observed in the human experiments. The conclusion is that grammatical cues such as word order are necessary to convey subjecthood and objecthood in a minority of naturally occurring transitive clauses; nevertheless, they can (a) provide an important source of redundancy and (b) are crucial for conveying intended meaning that cannot be inferred from the words alone, including descriptions of human interactions, where roles are often reversible (e.g., Ray helped Lu/Lu helped Ray), and expressing non-prototypical meanings (e.g., "The bone chewed the dog.").
研究动机与目标
- 确定语法线索(如词序、词形变化)在自然语言小句中用于识别主语的冗余频率。
- 评估仅依靠词义和合情合理性是否能可靠地完成主语识别,而无需依赖语法标记。
- 考察来自八个语系的30种语言中,语法线索冗余性的跨语言差异。
- 评估语法线索在非典型或可逆角色(例如‘Ray helped Lu’与‘Lu helped Ray’)中传达意义的作用。
提出的方法
- 对484名英语和俄语母语者进行了行为实验,呈现随机排列的主语-动词-宾语三词组,同时移除形态标记。
- 收集人类对每个三词组中哪个名词是动作主语的判断。
- 使用来自30种语言的多样化及物小句语料库,训练神经网络分类器完成相同任务。
- 测量分类器在不同语言中的表现,以评估在无语法线索条件下主语识别的准确性一致性。
- 使用跨语言语料库数据提取并分析在类型学上多样的语言中自然发生的及物小句。
- 比较人类与模型的表现,以评估仅靠意义和合情合理性在多大程度上可支持主语识别。
实验结果
研究问题
- RQ1在自然发生的及物小句中,语法线索用于主语识别的冗余频率如何?
- RQ2在不依赖语法标记的情况下,词义和合情合理性在多大程度上可独立决定主语身份?
- RQ3语法线索的冗余性在类型学上多样的语言中如何变化?
- RQ4在何种类型的小句中,语法线索是必不可少的,特别是在可逆角色或非典型语义的情况下?
- RQ5在无语法线索条件下,人类表现与神经网络表现相比如何?
主要发现
- 在英语中,参与者在移除形态和词序线索后,主语识别准确率达89%;在俄语中为87%。
- 神经网络分类器在30种语言中,仅依靠意义和合情合理性进行主语识别的中位准确率达87%。
- 在大多数简单及物小句中,词序和形态等语法线索具有冗余性,因为即使没有这些线索,主语识别依然高度准确。
- 在无语法线索条件下主语识别的高准确率表明,意义和合情合理性在大多数自然语境中已足够支持理解。
- 在非典型或可逆角色(如‘Ray helped Lu’与‘Lu helped Ray’)中,语法线索对于传达预期意义依然至关重要。
- 本研究证明了在不同语系中表现的一致性,表明主语识别的冗余性是一种跨语言普遍现象。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。