Open Access
REVIEW
Advancing Large Language Models for Low-Resource Languages: A Systematic Review of Pretraining, Adaptation, and Ethical Challenges
1 Department of Computer Science, George Mason University, Fairfax, VA, USA
2 Department of Computer Science, Colorado State University, Fort Collins, CO, USA
3 Faculty of Computer Science and Informatics, Berlin School of Business and Innovation, Berlin, Germany
4 Centre for Image and Vision Computing (CIVC), COE for Artificial Intelligence, Faculty of Artificial Intelligence and Engineering (FAIE), Multimedia University, Cyberjaya 63100, Selangor, Malaysia
5 Artificial Intelligence and Big Data Department, Endicott College, Woosong University, Daejeon, Republic of Korea
* Corresponding Authors: Jia Uddin. Email: ; Hezerul bin Abdul Karim. Email:
Computer Modeling in Engineering & Sciences 2026, 148(1), 4 https://doi.org/10.32604/cmes.2026.075507
Received 03 November 2025; Accepted 02 April 2026; Issue published 27 July 2026
Abstract
In recent years, the rapid advancement of Large Language Models (LLMs) has significantly transformed natural language processing (NLP), enabling impressive performance across a wide range of tasks. However, these developments have largely benefited high-resource languages, leaving many low-resource and underrepresented languages at risk of further digital marginalization. Addressing this imbalance is crucial to building more inclusive and culturally sustainable AI systems, which is motivating growing research interest in adapting LLMs for linguistically diverse and resource-scarce communities. This systematic review examines recent progress (2020–2025) in the pretraining and adaptation of LLMs for Low-Resource Languages (LRLs). Analysed 812 records obtained in the large databases and using PRISMA criteria, 140 core studies were identified. The innovations in data augmentation and parameter-efficient fine-tuning approaches can be outlined in this selection process. It combines major innovations on data-driven augmentation, parameter-efficient fine-tuning and morphologically rich and underrepresented language script-sensitive tokenization. The results highlight the growing effectiveness of culturally aware standards such as IrokoBench and BLEnD and show that approaches to lightweight adaptation eliminate high computational costs while maintaining language accuracy. The review focuses on the ethics in AI practice, the development of corpora through communities, and interdisciplinary research collaboration among computational linguists, social scientists, and digital humanists. The task of generating a diversified dataset, typology-conscious modelling strategies, and open-source multilingual benchmarks should be prioritized in future research as one possible solution to the existing digital language gap worldwide.Keywords
Supplementary Material
Supplementary Material FileCite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools