arrow
返回

Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability

delete2024-11-04
delete0
delete
OA
AI
T
Tyler A. Chang *
Z
Zhuowen Tu
B
Benjamin K. Bergen
DOI:10.1162/tacl_a_00708delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
How do language models learn to make predictions during pre-training? To study this, we extract learning curves from five autoregressive English language model pre-training runs, for 1M unseen tokens in context. We observe that the language models generate short repetitive phrases before learning to generate longer and more coherent text. We also find that individual tokens often exhibit sudden increases or decreases in loss that are surprisingly consistent across pre-training runs. To better understand these fluctuations, we quantify the final surprisal, within-run variability, age of acquisition, forgettability, and cross-run variability of learning curves for individual tokens in context. More frequent tokens reach lower final surprisals, exhibit less variability within and across pre-training runs, are learned earlier, and are less likely to be forgotten'' during pre-training. Higher n-gram probabilities further accentuate these effects. Independent of the target token, shorter and more frequent contexts correlate with marginally more stable and quickly acquired predictions. Based on our results, we argue for the existence of sequential learning dependencies between different model capabilities, and we characterize language model learning as early n-gram learning before gradual refinement of tail n-gram predictions.

期刊

T
Transactions of the Association for Computational Linguistics
IF:
6.9
论文数:
486
被引数:
5.7K

机构

University of California System 封面图
University of California System
学者数:
37.7W
论文数: 33.8W
被引数: 6.6K
引用论文

引用论文

An SL(2) Invariant Shape Median
err2010-02-25
err0
PREAI
errBenjamin Berkels; Gina Linkmann; Martin Rumpf
err分享
err收藏
Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond在实践中利用LLMs的力量: 关于ChatGPT及以后的调查
err2024-04-26
err137
errOAAI
errYang, Jingfeng; Jin, Hongye; Tang, Ruixiang; Han, Xiaotian; Feng, Qizhang; Jiang, Haoming; Zhong, Shaochen; Yin, Bing; Hu, Xia
err分享
err收藏
Hydrothermal Synthesis of Advanced Ceramic Powders
err2006-10-10
err0
PREAI
errWojciech L. Suchanek; Richard E. Riman
err分享
err收藏
Strong Prediction: Language Model Surprisal Explains Multiple N400 Effects强大的预测: 语言模型惊喜解释了多个N400效应
err2024-04-01
err13
errOAAI
errMichaelov, James A.; Bardolph, Megan D.; Van Petten, Cyma K.; Bergen, Benjamin K.; Coulson, Seana
err分享
err收藏
The influence of contextual diversity on word learning
err2015-11-23
err93
errOAAI
errJohns, Brendan T.; Dye, Melody; Jones, Michael N.
err分享
err收藏
学者 查看更多内容