返回
Complex Concept-Based Readability Estimation from Arabic Curriculum
DOI:10.1145/3770070.png)
摘要
En 中文
This article presents an approach to readability estimation that focuses on conceptual rather than linguistic complexity, using the extensive SaudiTextBooks textbooks. We introduce DARES 2.0, an enhanced concept-based readability training dataset designed to estimate the readability of Saudi educational texts. Building on DARES 1.0, DARES 2.0 extends the scope of conceptual complexity by replacing repetitive concepts and manually revising the input features with unique terms and their surrounding contexts from the SaudiTextBooks, spanning grades 1 to 12. The refined DARES 2.0 is employed to fine-tune pre-trained transformer models, including XLM-R Base, mBERT, AraELECTRA, AraBERTv2, and CAMeLBERTmix. The findings suggest that both the dataset and experimental setup require further development to ensure a larger, higher-quality dataset and to support more extensive fine-tuning experiments, in addition to exploring transfer learning from other languages and enhancing the diversity and richness of Arabic concepts. These developments pave the way for further advancements in concept-based readability estimation in educational contexts in future work.
Keyword:
Arabic text readability
conceptual complexity
readability estimation
Saudi curriculum
期刊
机构
引用论文
ReadNet: A Hierarchical Transformer Framework for Web Article Readability AnalysisReadNet:一种用于网页文章可读性分析的层次化Transformer框架

