arrow
返回

Deciphering Undersegmented Ancient Scripts Using Phonetic Prior

delete2021-02-01
delete10
delete
OA
AI
J
Jiaming Luo *
F
Frederik Hartmann
E
Enrico Santus
R
Regina Barzilay
Y
Yuan Cao
DOI:10.1162/tacl_a_00354delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Most undeciphered lost languages exhibit two characteristics that pose significant decipherment challenges: (1) the scripts are not fully segmented into words; (2) the closest known language is not determined. We propose a decipherment model that handles both of these challenges by building on rich linguistic constraints reflecting consistent patterns in historical sound change. We capture the natural phonological geometry by learning character embeddings based on the International Phonetic Alphabet (IPA). The resulting generative framework jointly models word segmentation and cognate alignment, informed by phonological constraints. We evaluate the model on both deciphered languages (Gothic, Ugaritic) and an undeciphered one (Iberian). The experiments show that incorporating phonetic geometry leads to clear and consistent gains. Additionally, we propose a measure for language closeness which correctly identifies related languages for Gothic and Ugaritic. For Iberian, the method does not show strong evidence supporting Basque as a related language, concurring with the favored position by the current scholarship.(1)

期刊

T
Transactions of the Association for Computational Linguistics
IF:
6.9
论文数:
486
被引数:
5.7K

机构

U
University of Konstanz
学者数:
6.1K
论文数: 5.1K
被引数: 7.7K
G
Google Incorporated
学者数:
3.5K
论文数: 1.8K
被引数: 8
引用论文

引用论文

Statehouse Democracy
err
IF0
err2010-08-04
err0
PREAI
errRobert S. Erikson; Gerald C. Wright; John P. McIver
err分享
err收藏
err分享
err收藏
没有更多内容