arrow
返回

A benchmark for evaluating Arabic contextualized word embedding models

delete2023-09-01
delete4
PRE
AI
A
Ashraf Elnagar *
S
Sane Yagi
Y
Youssef Mansour
L
Leena Lulu
S
Shehdeh Fareh
DOI:10.1016/j.ipm.2023.103452delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Word embeddings, which represent words as numerical vectors in a high-dimensional space, are contextualized by generating a unique vector representation for each sense of a word based on the surrounding words and sentence structure. They are typically generated using such deep learning models as BERT and trained on large amounts of text data and using self-supervised learning techniques. Resulting embeddings are highly effective at capturing the nuances of language, and have been shown to significantly improve the performance of numerous NLP tasks. Word embeddings represent textual records of human thinking, with all the mental relations that we utilize to produce the succession of sentences that make up texts and discourses. Consequently, the distributed representation of words within embeddings ought to capture the reasoning relations that hold texts together. This paper makes its contribution to the field by proposing a benchmark for the assessment of contextualized word embeddings that probes into their capability for true contextualization by inspecting how well they capture resemblance, contrariety, comparability, identity, relations in time and space, causation, analogy, and sense disambiguation. The proposed metrics adopt a triangulation approach, so they use (1) Hume's reasoning relations, (2) standard analogy, and (3) sense disambiguation. The benchmark has been evaluated against 22 Arabic contextualized embeddings and has proven to be capable of quantifying their differential performance in terms of these reasoning relations. Results of evaluation of the target embeddings revealed that they do take context into account and that they do reasonably well in sense disambiguation but have weakness in their identification of converseness, synonymy, complementarity, and analogy. Results also show that size of an embedding has diminishing returns because the highly frequent language patterns swamp low frequency patterns. Furthermore, the suggest that future research endeavors should not be concerned with the quantity of data as much as its quality, and that it should focus more on the representativeness of data, and on model architecture, design, and training.
Keyword:
Word contextualized embedding
Metrics
Intrinsic and extrinsic evaluation
Transformer
BERT

期刊

I
Information Processing and Management
IF:
6.9
论文数:
5.2K
被引数:
1.4W

机构

U
United Arab Emirates University
学者数:
8.8K
论文数: 7.4K
被引数: 10.0K
U
University of Sharjah
学者数:
5.9K
论文数: 5.5K
被引数: 8.8K
引用论文

引用论文

err分享
err收藏
A Panoramic Survey of Natural Language Processing in the Arab World
err2021-03-22
err35
errOAAI
errDarwish, Kareem; Habash, Nizar; Abbas, Mourad; Al-Khalifa, Hend; Al-Natsheh, Huseein T.; Bouamor, Houda; Bouzoubaa, Karim; Cavalli-Sforza, Violetta; El-Beltagy, Samhaa R.; El-Hajj, Wassim; Jarrar, Mustafa; Mubarak, Hamdy
err分享
err收藏
Arabic text classification using deep learning models
err2020-01-01
err137
PREAI
errElnagar, Ashraf; Al-Debsi, Ridhwan; Einea, Omar
err分享
err收藏
学者 查看更多内容