Data Augmentation Techniques for Arabic Short Text Processing

Comparative Insights and Future Directions

Authors

  • Ali Muftah Benomran Department of Computer Science, Faculty information technology, Alasmarya Islamic University, Zliten, Libya
  • Ashrf Ali Nasef Department of Computer Science, Faculty of Science, Alasmarya Islamic University, Zliten Libya
  • Aimen M.Rmis Department of Computer Science, Faculty of Science, Alasmarya Islamic University, Zliten, Libya
  • Salem Husein Almadhun Department of Computer, Faculty of Education, Elmergib University, Al Khums, Libya
  • Mamamer M Awinat Department of Computer Science, Faculty information technology, Alasmarya Islamic University, Zliten, Libya

DOI:

https://doi.org/10.59743/jaf.v10i1.960

Keywords:

Arabic NLP, Data Augmentation, Short Text Processing, Dialectal Variation, Low-Resource Languages, Comparative Review

Abstract

Arabic short texts pose challenges due to morphological complexity, dialectal diversity, and data scarcity. Data augmentation offers a solution, yet no review systematically compares techniques across tasks. We examine six strategies across five tasks using 22 studies (2019–2026). Results show no universal method works: synonym replacement boosts topic classification but harms NER; back-translation fails under dialectal shift; generative models improve sentiment analysis when filtered; feature-space resampling inflates metrics artificially. Three persistent issues cut across tasks: the diversity–fidelity trade-off, dialectal contamination, and computational cost. We propose a focused agenda emphasizing contrastive sampling, bias-aware generation, multimodal integration, and lightweight augmentation for low-resource dialects. This review provides actionable guidance for task-aware Arabic NLP frameworks.

Downloads

Download data is not yet available.

References

Abdhood, S. F., Omar, N., & Tiun, S. (2025). A Novel Data Augmentation Framework for Arabic Multi-Label Text Classification Using AraBART, AraGPT2, and Borderline-SMOTE. IEEE Access, 13, 169769–169778. https://doi.org/10.1109/ACCESS.2025.3609462

Alabdullah, A., Han, L., & Lin, C. (2025). Advancing Dialectal Arabic to Modern Standard Arabic Machine Translation. https://doi.org/10.21203/rs.3.rs-7510599/v1

Alayba, A. M. (2025). Arabic Natural Language Processing (NLP): A Comprehensive Review of Challenges, Techniques, and Emerging Trends. In Computers (Vol. 14, Number 11). Multidisciplinary Digital Publishing Institute (MDPI). https://doi.org/10.3390/computers14110497

Aljohani, E. (2024). Enhancing Arabic Fake News Detection: Evaluating Data Balancing Techniques Across Multiple Machine Learning Models. Engineering, Technology and Applied Science Research, 14(4), 15947–15956. https://doi.org/10.48084/etasr.8019

Al-Khazaleh, M. J., Alian, M., & Jaradat, M. A. (2024). Sentiment analysis of imbalanced Arabic data using sampling techniques and classification algorithms. Bulletin of Electrical Engineering and Informatics, 13(1), 607–618. https://doi.org/10.11591/eei.v13i1.5886

Almutairi, S., & Alotaibi, F. (2023). A Comparative Analysis for Arabic Sentiment Analysis Models In E-Marketing Using Deep Learning Techniques. In Journal of Engineering and Applied Sciences (Vol. 10, Number 1).

Alotaibi, A., & Nadeem, F. (2024). Leveraging Social Media and Deep Learning for Sentiment Analysis for Smart Governance: A Case Study of Public Reactions to Educational Reforms in Saudi Arabia. Computers, 13(11). https://doi.org/10.3390/computers13110280

Alrashidi, B., Jamal, A., & Alkhathlan, A. (2023). Abusive Content Detection in Arabic Tweets Using Multi-Task Learning and Transformer-Based Models. Applied Sciences (Switzerland), 13(10). https://doi.org/10.3390/app13105825

Badri, N., Kboubi, F., & Habacha Chaibi, A. (2024). Abusive and Hate speech Classification in Arabic Text Using Pre-trained Language Models and Data Augmentation. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(11). https://doi.org/10.1145/3679049

Bayer, M., Kaufhold, M. A., & Reuter, C. (2022). A Survey on Data Augmentation for Text Classification. ACM Computing Surveys, 55(7). https://doi.org/10.1145/3544558

Chen, S., Zhang, Y., & Yang, Q. (2024). Multi-Task Learning in Natural Language Processing: An Overview. ACM Computing Surveys, 56(12). https://doi.org/10.1145/3663363

Cooper, R., Kliesner, K. W., & Zenker, S. (2024). Contrastive Meta-Learner for Automatic Text Labeling and Semantic Textual Similarity. IEEE Access, 12, 166792–166799. https://doi.org/10.1109/ACCESS.2024.3424401

Dahou, A. H., Cheragui, M. A., Abdedaiem, A., & Mathiak, B. (2024). Enhancing Model Performance through Translation-based Data Augmentation in the context of Fake News Detection. Procedia Computer Science, 244, 342–352. https://doi.org/10.1016/j.procs.2024.10.208

Elgobshawi, A. E. (2024). Conceptualization of Morphological Roots in Arabic and English: A Contrastive Analysis. World Journal of English Language, 14(5), 436–443. https://doi.org/10.5430/wjel.v14n5p436

Elkhbir, N. (2024). Information extraction for arabic and its dialects [Thesis, UNIVERSITÉ SORBONNE PARIS NORD]. https://doi.org/10.70675/702bb6e7z1efcz4563z9539zb485ffc11160

Elnaka, A., Nael, O., Afifi, H., & Sharaf, N. (2021). AraScore: Investigating Response-Based Arabic Short Answer Scoring. Procedia CIRP, 189, 282–291. https://doi.org/10.1016/j.procs.2021.05.091

ElSabagh, A. A., Azab, S. S., & Hefny, H. A. (2025). A comprehensive survey on Arabic text augmentation: approaches, challenges, and applications. In Neural Computing and Applications. Springer Science and Business Media Deutschland GmbH. https://doi.org/10.1007/s00521-025-11020-z

Feng, S. Y., Gangal, V., Wei, J., Chandar, S., Vosoughi, S., Mitamura, T., & Hovy, E. (2021). A Survey of Data Augmentation Approaches for NLP. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 968–988. https://doi.org/10.18653/v1/2021.findings-acl.84

Gagliardi, I., & Artese, M. T. (2023). Ensemble-Based Short Text Similarity: An Easy Approach for Multilingual Datasets Using Transformers and WordNet in Real-World Scenarios. Big Data and Cognitive Computing, 7(4). https://doi.org/10.3390/bdcc7040158

Ghazoui, B., Bazi, I. El, Essadik, I., Benali, B. A., & Moussa, H. (2026). Robust Arabic tweet NER via label-aware data augmentation and AraBERTv2. Bulletin of Electrical Engineering and Informatics, 15(1), 740–754. https://doi.org/10.11591/eei.v15i1.10462

Habbat, N., Nouri, H., Anoun, H., & Hassouni, L. (2023). Using AraGPT and ensemble deep learning model for sentiment analysis on Arabic imbalanced dataset. ITM Web of Conferences, 52, 02008. https://doi.org/10.1051/itmconf/20235202008

Habberrih, A., & Abuzaraida, M. A. (2024). A review of the available Arabic dialects datasets for Sentiment Analysis. Journal of Sustainable Research in Applied Sciences, 30(2), 30–37.

Hossain, M. M., Hossain, M. S., Safran, M., Alfarhood, S., Alfarhood, M., & Mridha, M. F. (2024). A Hybrid Attention-Based Transformer Model for Arabic News Classification Using Text Embedding and Deep Learning. IEEE Access, 12, 198046–198066. https://doi.org/10.1109/ACCESS.2024.3522061

Hsu, T.-W., Chen, C.-C., Huang, H.-H., & Chen, H.-H. (2021). Semantics-Preserved Data Augmentation for Aspect-Based Sentiment Analysis. Conference on Empirical Methods in Natural Language Processing, 4417–4422.

Lee, J.-M., & Ha, T.-B. (2023). Unsupervised Text Embedding Space Generation Using Generative Adversarial Networks for Text Synthesis. https://doi.org/10.3384/nejlt.2000-1533.2023.4855

Li, B., Hou, Y., & Che, W. (2022). Data augmentation approaches in natural language processing: A survey. AI Open, 3, 71–90. https://doi.org/10.1016/j.aiopen.2022.03.001

Mohamed, E. A., Ismail, W. N., Ibrahim, O. A. S., & Younis, E. M. G. (2024). A two-stage framework for Arabic social media text misinformation detection combining data augmentation and AraBERT. Social Network Analysis and Mining, 14(1). https://doi.org/10.1007/s13278-024-01201-4

Mumuni, A., & Mumuni, F. (2022). Data augmentation: A comprehensive survey of modern approaches. In Array (Vol. 16). Elsevier B.V. https://doi.org/10.1016/j.array.2022.100258

Nacar, O., Sibaee, S., Ahmed, S., Alharbi, A. I., Ghouti, L., & Koubaa, A. (2024). ASOS at OSACT6 Shared Task: Investigation of Data Augmentation in Arabic Dialect-MSA Translation. ASOS at OSACT6 Shared Task: Investigation of Data Augmentation in Arabic Dialect-MSA Translation, 104–111.

Refai, D., Abu-Soud, S., & Abdel-Rahman, M. J. (2023a). Data Augmentation Using Transformers and Similarity Measures for Improving Arabic Text Classification. IEEE Access, 11, 132516–132531. https://doi.org/10.1109/ACCESS.2023.3336311

Refai, D., Abu-Soud, S., & Abdel-Rahman, M. J. (2023b). Data Augmentation Using Transformers and Similarity Measures for Improving Arabic Text Classification. IEEE Access, 11, 132516–132531. https://doi.org/10.1109/ACCESS.2023.3336311

Sabty, C., Omar, I., Wasfalla, F., Islam, M., & Abdennadher, S. (2021). Data Augmentation Techniques on Arabic Data for Named Entity Recognition. Procedia Computer Science, 189, 292–299. https://doi.org/10.1016/j.procs.2021.05.092

Talafha, B., Fadel, A., Al-Ayyoub, M., Jararweh, Y., Al-Smadi, M., & Juola, P. (2019). Team JUST at the MADAR Shared Task on Arabic Fine-Grained Dialect Identification. Proceedings of the Fourth Arabic Natural Language Processing, 285–289. https://doi.org/10.18653/v1/W19-4638

Worth, P. J. (2023). Word Embeddings and Semantic Spaces in Natural Language Processing. International Journal of Intelligence Science, 13(01), 1–21. https://doi.org/10.4236/ijis.2023.131001

Downloads

Published

30-06-2026

How to Cite

Benomran, A. M. ., Nasef, A. A. ., Rmis, A. M., Almadhun, S. H. ., & Awinat, M. M. . (2026). Data Augmentation Techniques for Arabic Short Text Processing: Comparative Insights and Future Directions. Journal of the Academic Forum, 10(1), 153-179. https://doi.org/10.59743/jaf.v10i1.960