[论文解读] Frequency Explains the Inverse Correlation of Large Language Models' Size, Training Data Amount, and Surprisal's Fit to Reading Times
本文識別出詞頻是解釋為何更大規模語言模型(LLMs)及訓練數據更多之模型,其 surprisal 評估與人類閱讀時間之間的擬合度下降之關鍵因素。透過對四種模型家族與語料的分析,本文顯示較大模型會因對罕見詞學習過於複雜的關聯而過度擬合,進而導致對人類閱讀行為的預測能力下降,特別是在低頻詞上表現更差。
Recent studies have shown that as Transformer-based language models become larger and are trained on very large amounts of data, the fit of their surprisal estimates to naturalistic human reading times degrades. The current work presents a series of analyses showing that word frequency is a key explanatory factor underlying these two trends. First, residual errors from four language model families on four corpora show that the inverse correlation between model size and fit to reading times is the strongest on the subset of least frequent words, which is driven by excessively accurate predictions of larger model variants. Additionally, training dynamics reveal that during later training steps, all model variants learn to predict rare words and that larger model variants do so more accurately, which explains the detrimental effect of both training data amount and model size on fit to reading times. Finally, a feature attribution analysis demonstrates that larger model variants are able to accurately predict rare words based on both an effectively longer context window size as well as stronger local associations compared to smaller model variants. Taken together, these results indicate that Transformer-based language models' surprisal estimates diverge from human-like expectations due to the superhumanly complex associations they learn for predicting rare words.
研究动机与目标
- 探討為何更大規模語言模型及訓練數據更多的模型,其 surprisal 評估與人類閱讀時間之間的擬合度逐漸下降。
- 確認詞頻是否調節模型規模/訓練數據與 surprisal 預測能力之間的反向相關性。
- 檢視訓練動態與模型架構如何影響對罕見詞與常見詞的預測。
- 評估大規模模型對罕見詞過度精確的預測是否削弱其模擬人類處理難度的能力。
- 探討在自然語境下的閱讀理解中,詞頻效應與可預測性效應是否可分離。
提出的方法
- 在四個自然語境下的閱讀時間語料(Natural Stories、Dundee、GECO、Provo)上,評估四種基於 Transformer 的語言模型家族(如 GPT、BERT 變體)的 surprisal 評估。
- 依詞頻分位數,計算模型預測 surprisal 與實際閱讀時間之間的殘差誤差。
- 追蹤不同模型變體的訓練動態,以分析在增加訓練 token 數量的過程中,罕見詞何時及如何被學習。
- 應用特徵歸因技術(如基於梯度或注意力遮蔽)以分離上下文視窗長度與局部注意力強度對罕見詞預測的貢獻。
- 使用迴歸模型評估 surprisal 評估與閱讀時間的擬合度,並在有無詞頻分層的情況下進行比較。
- 比較不同規模(小、中、大)的模型變體,並在逐漸增加的訓練數據量(最多 2B tokens)下訓練,以區分規模與資料量的影響。

实验结果
研究问题
- RQ1模型規模與 surprisal 對閱讀時間擬合度之間的反向相關性是否因詞頻而異?
- RQ2訓練動態如何影響不同規模模型變體對罕見詞預測的學習?
- RQ3與小規模模型相比,大規模模型在多大程度上學會了更精確的罕見詞表徵?
- RQ4哪些機制(如上下文視窗長度、注意力強度)導致大規模模型在罕見詞預測上表現更優?
- RQ5語言模型的 surprisal 與人類閱讀時間之間的分歧,是否可由對罕見詞中與詞頻無關的複雜關聯過度擬合來解釋?
主要发现
- 模型規模與對閱讀時間擬合度之間的反向相關性,在最不常見的詞上最為顯著,其中大規模模型產生的 surprisal 評估顯著更精確。
- 大規模模型變體在預測罕見詞方面比小規模模型更為精確,特別是在訓練後期階段,這導致其與人類閱讀時間的擬合度下降。
- 訓練動態顯示,所有模型變體僅在經歷大量訓練資料(例如超過 2B tokens)後才開始學習罕見詞預測,且大規模模型能更快達成更高準確度。
- 特徵歸因分析顯示,大規模模型能更精確預測罕見詞,是因為其具有更長的有效上下文視窗與更強的局部注意力關聯。
- 對罕見詞的過度精確預測——由超人類級複雜關聯驅動——解釋了為何大規模模型的 surprisal 與人類處理預期產生分歧。
- 限制上下文視窗大小或減弱局部注意力關聯,可改善 surprisal 對閱讀時間的擬合度,確認對罕見詞模式的過度擬合會損害對人類處理過程的預測。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。