arrow
Return

Effective low-resource Arabic dialect identification

delete2026-08-22
delete0
delete
OA
AI
M
Mohammed Abdelmajeed
Z
Zheng Jiangbin
K
Khidir Shaib Mohamed *
A
Abdalilah Alhalangy
S
Suhail Abdullah Alsaqer
S
Sofian A.A. Saad
DOI:10.1016/j.rineng.2026.112558delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Arabic Dialect Identification (ADI) remains challenging because of scarce labeled resources and substantial variation between Modern Standard Arabic (MSA) and regional dialects. While existing approaches predominantly rely on fine-tuning pre-trained language models (PLMs), their performance degrades sharply in low-resource scenarios with only a few labeled examples per dialect. To address this limitation, we propose an adversarial semi-supervised learning framework that leverages abundant unlabeled dialectal text to improve robustness and generalization under data scarcity. The proposed method integrates a Gradient Reversal Layer (GRL) to encourage source-invariant representation learning, guiding the PLM toward dialect-discriminative yet source-stable features. The framework jointly optimizes supervised dialect classification and adversarial source alignment through a composite objective. Specifically, the effective encoder objective becomes L f = L d − λ L s through gradient reversal, where λ controls the strength of adversarial alignment. We evaluate the approach under 5-, 8-, and 10-shot settings on MADAR-2, MADAR-6, MADAR-9, MADAR-26, and NADI. Extensive experiments demonstrate that the proposed method improves upon standard fine-tuning in most evaluated configurations, with particularly notable gains in accuracy and macro-F1 under ultra-low-resource conditions. The proposed framework also tends to reduce performance variability across repeated few-shot runs, although the magnitude of the improvement depends on the backbone PLM and dataset. These results indicate that adversarial alignment provides an effective strategy for mitigating data scarcity in Arabic Dialect Identification. Although the framework is based on general domain-adaptation principles, its applicability beyond Arabic dialect identification requires further empirical validation.. The implementation is publicly available at https://github.com/amurtadha/ADI-main/ .
Keywords:
Arabic dialect identification
Natural language processing
Bidirectional encoder representations from transformers
Pre-trained language models
Gradient reversal layer

Journal

Results in Engineering cover
Results in Engineering
IF:
7.9
Papers:
1.1W
Citations:
1.7W

Organization

N
northwestern polytechnical university
Scholars:
1.2W
Papers: 4.4K
Citations: 0
Q
Qassim University
Scholars:
5.6K
Papers: 5.4K
Citations: 5.0K