Return
Colonialism in code: linguistic hegemony and epistemic exclusion in large language models
L
DOI:10.1007/s00146-026-03184-6.png)
Abstract
En 中文
This review examines how large language models (LLMs) reflect and reinforce global patterns of epistemic inequality through linguistic hegemony and the exclusion of marginalized knowledge systems. While often framed as neutral tools, LLMs embed epistemological assumptions drawn from dominant cultures and worldviews, particularly those encoded in high-resourced languages, such as English. Drawing on postcolonial theory and the concept of epistemic injustice, this article synthesizes current research and debates to trace parallels between algorithmic language modeling and historical practices of linguistic domination and knowledge extraction. Through a critical discourse analysis of ChatGPT-4o and DeepSeek-V3’s responses to culturally sensitive prompts in both English and Chinese, this review illustrates how both models converge around liberal, Anglophone norms, even when operating in multilingual contexts. These findings are situated within a broader review of AI policy frameworks and technical efforts to support language inclusion. This review argues that without community-led interventions and structural change, LLMs risk entrenching algorithmic forms of cultural imperialism. To support linguistic diversity and global justice, AI development must move beyond extractive data practices and engage with plural, situated knowledge systems—including Indigenous epistemologies. This review contributes to emerging efforts to decolonize AI by foregrounding the cultural politics of language and knowledge representation in machine learning systems.
Keywords:
Linguistic hegemony
Epistemic injustice
Large language models
Decolonizing AI
Knowledge systems
Critical discourse analysis
Journal
A
IF:
4.7
Papers:
405
Citations:
0
