Data Augmentation Techniques for Arabic Short Text Processing

Comparative Insights and Future Directions

Authors

  • Ali Muftah Benomran Department of Computer Science, Faculty information technology, Alasmarya Islamic University, Zliten, Libya
  • Ashrf Ali Nasef Department of Computer Science, Faculty of Science, Alasmarya Islamic University, Zliten Libya
  • Aimen M.Rmis Department of Computer Science, Faculty of Science, Alasmarya Islamic University, Zliten, Libya
  • Salem Husein Almadhun Department of Computer, Faculty of Education, Elmergib University, Al Khums, Libya
  • Mamamer M Awinat Department of Computer Science, Faculty information technology, Alasmarya Islamic University, Zliten, Libya

DOI:

https://doi.org/10.59743/jaf.v10i1.960

Keywords:

Arabic NLP, Data Augmentation, Short Text Processing, Dialectal Variation, Low-Resource Languages, Comparative Review

Abstract

Arabic short texts pose challenges due to morphological complexity, dialectal diversity, and data scarcity. Data augmentation offers a solution, yet no review systematically compares techniques across tasks. We examine six strategies across five tasks using 22 studies (2019–2026). Results show no universal method works: synonym replacement boosts topic classification but harms NER; back-translation fails under dialectal shift; generative models improve sentiment analysis when filtered; feature-space resampling inflates metrics artificially. Three persistent issues cut across tasks: the diversity–fidelity trade-off, dialectal contamination, and computational cost. We propose a focused agenda emphasizing contrastive sampling, bias-aware generation, multimodal integration, and lightweight augmentation for low-resource dialects. This review provides actionable guidance for task-aware Arabic NLP frameworks.

Downloads

Download data is not yet available.

References

Abdhood, S. F., Omar, N., & Tiun, S. (2025). A Novel Data Augmentation Framework for Arabic Multi-Label Text Classification Using AraBART, AraGPT2, and Borderline-SMOTE. IEEE Access, 13, 169769–169778. https://doi.org/10.1109/ACCESS.2025.3609462 DOI: https://doi.org/10.1109/ACCESS.2025.3609462

Alabdullah, A., Han, L., & Lin, C. (2025). Advancing Dialectal Arabic to Modern Standard Arabic Machine Translation. https://doi.org/10.21203/rs.3.rs-7510599/v1 DOI: https://doi.org/10.21203/rs.3.rs-7510599/v1

Alayba, A. M. (2025). Arabic Natural Language Processing (NLP): A Comprehensive Review of Challenges, Techniques, and Emerging Trends. In Computers (Vol. 14, Number 11). Multidisciplinary Digital Publishing Institute (MDPI). https://doi.org/10.3390/computers14110497 DOI: https://doi.org/10.3390/computers14110497

Aljohani, E. (2024). Enhancing Arabic Fake News Detection: Evaluating Data Balancing Techniques Across Multiple Machine Learning Models. Engineering, Technology and Applied Science Research, 14(4), 15947–15956. https://doi.org/10.48084/etasr.8019 DOI: https://doi.org/10.48084/etasr.8019

Al-Khazaleh, M. J., Alian, M., & Jaradat, M. A. (2024). Sentiment analysis of imbalanced Arabic data using sampling techniques and classification algorithms. Bulletin of Electrical Engineering and Informatics, 13(1), 607–618. https://doi.org/10.11591/eei.v13i1.5886 DOI: https://doi.org/10.11591/eei.v13i1.5886

Almutairi, S., & Alotaibi, F. (2023). A Comparative Analysis for Arabic Sentiment Analysis Models In E-Marketing Using Deep Learning Techniques. In Journal of Engineering and Applied Sciences (Vol. 10, Number 1). DOI: https://doi.org/10.5455/jeas.2023050102

Alotaibi, A., & Nadeem, F. (2024). Leveraging Social Media and Deep Learning for Sentiment Analysis for Smart Governance: A Case Study of Public Reactions to Educational Reforms in Saudi Arabia. Computers, 13(11). https://doi.org/10.3390/computers13110280 DOI: https://doi.org/10.3390/computers13110280

Alrashidi, B., Jamal, A., & Alkhathlan, A. (2023). Abusive Content Detection in Arabic Tweets Using Multi-Task Learning and Transformer-Based Models. Applied Sciences (Switzerland), 13(10). https://doi.org/10.3390/app13105825 DOI: https://doi.org/10.3390/app13105825

Badri, N., Kboubi, F., & Habacha Chaibi, A. (2024). Abusive and Hate speech Classification in Arabic Text Using Pre-trained Language Models and Data Augmentation. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(11). https://doi.org/10.1145/3679049 DOI: https://doi.org/10.1145/3679049

Bayer, M., Kaufhold, M. A., & Reuter, C. (2022). A Survey on Data Augmentation for Text Classification. ACM Computing Surveys, 55(7). https://doi.org/10.1145/3544558 DOI: https://doi.org/10.1145/3544558

Chen, S., Zhang, Y., & Yang, Q. (2024). Multi-Task Learning in Natural Language Processing: An Overview. ACM Computing Surveys, 56(12). https://doi.org/10.1145/3663363 DOI: https://doi.org/10.1145/3663363

Cooper, R., Kliesner, K. W., & Zenker, S. (2024). Contrastive Meta-Learner for Automatic Text Labeling and Semantic Textual Similarity. IEEE Access, 12, 166792–166799. https://doi.org/10.1109/ACCESS.2024.3424401 DOI: https://doi.org/10.1109/ACCESS.2024.3424401

Dahou, A. H., Cheragui, M. A., Abdedaiem, A., & Mathiak, B. (2024). Enhancing Model Performance through Translation-based Data Augmentation in the context of Fake News Detection. Procedia Computer Science, 244, 342–352. https://doi.org/10.1016/j.procs.2024.10.208 DOI: https://doi.org/10.1016/j.procs.2024.10.208

Elgobshawi, A. E. (2024). Conceptualization of Morphological Roots in Arabic and English: A Contrastive Analysis. World Journal of English Language, 14(5), 436–443. https://doi.org/10.5430/wjel.v14n5p436 DOI: https://doi.org/10.5430/wjel.v14n5p436

Elkhbir, N. (2024). Information extraction for arabic and its dialects [Thesis, UNIVERSITÉ SORBONNE PARIS NORD]. https://doi.org/10.70675/702bb6e7z1efcz4563z9539zb485ffc11160 DOI: https://doi.org/10.70675/702bb6e7z1efcz4563z9539zb485ffc11160

Elnaka, A., Nael, O., Afifi, H., & Sharaf, N. (2021). AraScore: Investigating Response-Based Arabic Short Answer Scoring. Procedia CIRP, 189, 282–291. https://doi.org/10.1016/j.procs.2021.05.091 DOI: https://doi.org/10.1016/j.procs.2021.05.091

ElSabagh, A. A., Azab, S. S., & Hefny, H. A. (2025). A comprehensive survey on Arabic text augmentation: approaches, challenges, and applications. In Neural Computing and Applications. Springer Science and Business Media Deutschland GmbH. https://doi.org/10.1007/s00521-025-11020-z DOI: https://doi.org/10.1007/s00521-025-11020-z

Feng, S. Y., Gangal, V., Wei, J., Chandar, S., Vosoughi, S., Mitamura, T., & Hovy, E. (2021). A Survey of Data Augmentation Approaches for NLP. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 968–988. https://doi.org/10.18653/v1/2021.findings-acl.84 DOI: https://doi.org/10.18653/v1/2021.findings-acl.84

Gagliardi, I., & Artese, M. T. (2023). Ensemble-Based Short Text Similarity: An Easy Approach for Multilingual Datasets Using Transformers and WordNet in Real-World Scenarios. Big Data and Cognitive Computing, 7(4). https://doi.org/10.3390/bdcc7040158 DOI: https://doi.org/10.3390/bdcc7040158

Ghazoui, B., Bazi, I. El, Essadik, I., Benali, B. A., & Moussa, H. (2026). Robust Arabic tweet NER via label-aware data augmentation and AraBERTv2. Bulletin of Electrical Engineering and Informatics, 15(1), 740–754. https://doi.org/10.11591/eei.v15i1.10462 DOI: https://doi.org/10.11591/eei.v15i1.10462

Habbat, N., Nouri, H., Anoun, H., & Hassouni, L. (2023). Using AraGPT and ensemble deep learning model for sentiment analysis on Arabic imbalanced dataset. ITM Web of Conferences, 52, 02008. https://doi.org/10.1051/itmconf/20235202008 DOI: https://doi.org/10.1051/itmconf/20235202008

Habberrih, A., & Abuzaraida, M. A. (2024). A review of the available Arabic dialects datasets for Sentiment Analysis. Journal of Sustainable Research in Applied Sciences, 30(2), 30–37. DOI: https://doi.org/10.36602/jsras.2024.1.2.1

Hossain, M. M., Hossain, M. S., Safran, M., Alfarhood, S., Alfarhood, M., & Mridha, M. F. (2024). A Hybrid Attention-Based Transformer Model for Arabic News Classification Using Text Embedding and Deep Learning. IEEE Access, 12, 198046–198066. https://doi.org/10.1109/ACCESS.2024.3522061 DOI: https://doi.org/10.1109/ACCESS.2024.3522061

Hsu, T.-W., Chen, C.-C., Huang, H.-H., & Chen, H.-H. (2021). Semantics-Preserved Data Augmentation for Aspect-Based Sentiment Analysis. Conference on Empirical Methods in Natural Language Processing, 4417–4422. DOI: https://doi.org/10.18653/v1/2021.emnlp-main.362

Lee, J.-M., & Ha, T.-B. (2023). Unsupervised Text Embedding Space Generation Using Generative Adversarial Networks for Text Synthesis. https://doi.org/10.3384/nejlt.2000-1533.2023.4855 DOI: https://doi.org/10.3384/nejlt.2000-1533.2023.4855

Li, B., Hou, Y., & Che, W. (2022). Data augmentation approaches in natural language processing: A survey. AI Open, 3, 71–90. https://doi.org/10.1016/j.aiopen.2022.03.001 DOI: https://doi.org/10.1016/j.aiopen.2022.03.001

Mohamed, E. A., Ismail, W. N., Ibrahim, O. A. S., & Younis, E. M. G. (2024). A two-stage framework for Arabic social media text misinformation detection combining data augmentation and AraBERT. Social Network Analysis and Mining, 14(1). https://doi.org/10.1007/s13278-024-01201-4 DOI: https://doi.org/10.1007/s13278-024-01201-4

Mumuni, A., & Mumuni, F. (2022). Data augmentation: A comprehensive survey of modern approaches. In Array (Vol. 16). Elsevier B.V. https://doi.org/10.1016/j.array.2022.100258 DOI: https://doi.org/10.1016/j.array.2022.100258

Nacar, O., Sibaee, S., Ahmed, S., Alharbi, A. I., Ghouti, L., & Koubaa, A. (2024). ASOS at OSACT6 Shared Task: Investigation of Data Augmentation in Arabic Dialect-MSA Translation. ASOS at OSACT6 Shared Task: Investigation of Data Augmentation in Arabic Dialect-MSA Translation, 104–111. DOI: https://doi.org/10.63317/34u3y58xvz4e

Refai, D., Abu-Soud, S., & Abdel-Rahman, M. J. (2023a). Data Augmentation Using Transformers and Similarity Measures for Improving Arabic Text Classification. IEEE Access, 11, 132516–132531. https://doi.org/10.1109/ACCESS.2023.3336311

Refai, D., Abu-Soud, S., & Abdel-Rahman, M. J. (2023b). Data Augmentation Using Transformers and Similarity Measures for Improving Arabic Text Classification. IEEE Access, 11, 132516–132531. https://doi.org/10.1109/ACCESS.2023.3336311 DOI: https://doi.org/10.1109/ACCESS.2023.3336311

Sabty, C., Omar, I., Wasfalla, F., Islam, M., & Abdennadher, S. (2021). Data Augmentation Techniques on Arabic Data for Named Entity Recognition. Procedia Computer Science, 189, 292–299. https://doi.org/10.1016/j.procs.2021.05.092 DOI: https://doi.org/10.1016/j.procs.2021.05.092

Talafha, B., Fadel, A., Al-Ayyoub, M., Jararweh, Y., Al-Smadi, M., & Juola, P. (2019). Team JUST at the MADAR Shared Task on Arabic Fine-Grained Dialect Identification. Proceedings of the Fourth Arabic Natural Language Processing, 285–289. https://doi.org/10.18653/v1/W19-4638 DOI: https://doi.org/10.18653/v1/W19-4638

Worth, P. J. (2023). Word Embeddings and Semantic Spaces in Natural Language Processing. International Journal of Intelligence Science, 13(01), 1–21. https://doi.org/10.4236/ijis.2023.131001 DOI: https://doi.org/10.4236/ijis.2023.131001

Downloads

Published

30-06-2026

How to Cite

Benomran, A. M. ., Nasef, A. A. ., Rmis, A. M., Almadhun, S. H. ., & Awinat, M. M. . (2026). Data Augmentation Techniques for Arabic Short Text Processing: Comparative Insights and Future Directions. Journal of the Academic Forum, 10(1), 153-179. https://doi.org/10.59743/jaf.v10i1.960