返回
Arithmetic with language models: From memorization to computation
DOI:10.1016/j.neunet.2024.106550.png)
摘要
En 中文
A better understanding of the emergent computation and problem-solving capabilities of recent large language models is of paramount importance to further improve them and broaden their applicability. This work investigates how a language model, trained to predict the next token, can perform arithmetic computations generalizing beyond training data. Binary addition and multiplication constitute a good testbed for this purpose, since they require a very small vocabulary and exhibit relevant input/output discontinuities making smooth input interpolation ineffective for novel data. We successfully trained a light language model to learn these tasks and ran a number of experiments to investigate the extrapolation capabilities and internal information processing. Our findings support the hypothesis that the language model works as an Encoding-Regression- Decoding machine where the computation takes place in the value space once the input token representation is mapped to an appropriate internal representation.
Keyword:
Language models
AI explainability
Probing
Interpretability
Arithmetic
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.3
论文数:
7.8K
被引数:
3.0W
机构
引用论文
没有更多内容

