Data Augmentation Techniques for Arabic Short Text Processing
Comparative Insights and Future Directions
DOI:
https://doi.org/10.59743/jaf.v10i1.960Keywords:
Arabic NLP, Data Augmentation, Short Text Processing, Dialectal Variation, Low-Resource Languages, Comparative ReviewAbstract
Arabic short texts pose challenges due to morphological complexity, dialectal diversity, and data scarcity. Data augmentation offers a solution, yet no review systematically compares techniques across tasks. We examine six strategies across five tasks using 22 studies (2019–2026). Results show no universal method works: synonym replacement boosts topic classification but harms NER; back-translation fails under dialectal shift; generative models improve sentiment analysis when filtered; feature-space resampling inflates metrics artificially. Three persistent issues cut across tasks: the diversity–fidelity trade-off, dialectal contamination, and computational cost. We propose a focused agenda emphasizing contrastive sampling, bias-aware generation, multimodal integration, and lightweight augmentation for low-resource dialects. This review provides actionable guidance for task-aware Arabic NLP frameworks.
Downloads
References
Abdhood, S. F., Omar, N., & Tiun, S. (2025). A Novel Data Augmentation Framework for Arabic Multi-Label Text Classification Using AraBART, AraGPT2, and Borderline-SMOTE. IEEE Access, 13, 169769–169778. https://doi.org/10.1109/ACCESS.2025.3609462 DOI: https://doi.org/10.1109/ACCESS.2025.3609462
Alabdullah, A., Han, L., & Lin, C. (2025). Advancing Dialectal Arabic to Modern Standard Arabic Machine Translation. https://doi.org/10.21203/rs.3.rs-7510599/v1 DOI: https://doi.org/10.21203/rs.3.rs-7510599/v1
Alayba, A. M. (2025). Arabic Natural Language Processing (NLP): A Comprehensive Review of Challenges, Techniques, and Emerging Trends. In Computers (Vol. 14, Number 11). Multidisciplinary Digital Publishing Institute (MDPI). https://doi.org/10.3390/computers14110497 DOI: https://doi.org/10.3390/computers14110497
Aljohani, E. (2024). Enhancing Arabic Fake News Detection: Evaluating Data Balancing Techniques Across Multiple Machine Learning Models. Engineering, Technology and Applied Science Research, 14(4), 15947–15956. https://doi.org/10.48084/etasr.8019 DOI: https://doi.org/10.48084/etasr.8019
Al-Khazaleh, M. J., Alian, M., & Jaradat, M. A. (2024). Sentiment analysis of imbalanced Arabic data using sampling techniques and classification algorithms. Bulletin of Electrical Engineering and Informatics, 13(1), 607–618. https://doi.org/10.11591/eei.v13i1.5886 DOI: https://doi.org/10.11591/eei.v13i1.5886
Almutairi, S., & Alotaibi, F. (2023). A Comparative Analysis for Arabic Sentiment Analysis Models In E-Marketing Using Deep Learning Techniques. In Journal of Engineering and Applied Sciences (Vol. 10, Number 1). DOI: https://doi.org/10.5455/jeas.2023050102
Alotaibi, A., & Nadeem, F. (2024). Leveraging Social Media and Deep Learning for Sentiment Analysis for Smart Governance: A Case Study of Public Reactions to Educational Reforms in Saudi Arabia. Computers, 13(11). https://doi.org/10.3390/computers13110280 DOI: https://doi.org/10.3390/computers13110280
Alrashidi, B., Jamal, A., & Alkhathlan, A. (2023). Abusive Content Detection in Arabic Tweets Using Multi-Task Learning and Transformer-Based Models. Applied Sciences (Switzerland), 13(10). https://doi.org/10.3390/app13105825 DOI: https://doi.org/10.3390/app13105825
Badri, N., Kboubi, F., & Habacha Chaibi, A. (2024). Abusive and Hate speech Classification in Arabic Text Using Pre-trained Language Models and Data Augmentation. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(11). https://doi.org/10.1145/3679049 DOI: https://doi.org/10.1145/3679049
Bayer, M., Kaufhold, M. A., & Reuter, C. (2022). A Survey on Data Augmentation for Text Classification. ACM Computing Surveys, 55(7). https://doi.org/10.1145/3544558 DOI: https://doi.org/10.1145/3544558
Chen, S., Zhang, Y., & Yang, Q. (2024). Multi-Task Learning in Natural Language Processing: An Overview. ACM Computing Surveys, 56(12). https://doi.org/10.1145/3663363 DOI: https://doi.org/10.1145/3663363
Cooper, R., Kliesner, K. W., & Zenker, S. (2024). Contrastive Meta-Learner for Automatic Text Labeling and Semantic Textual Similarity. IEEE Access, 12, 166792–166799. https://doi.org/10.1109/ACCESS.2024.3424401 DOI: https://doi.org/10.1109/ACCESS.2024.3424401
Dahou, A. H., Cheragui, M. A., Abdedaiem, A., & Mathiak, B. (2024). Enhancing Model Performance through Translation-based Data Augmentation in the context of Fake News Detection. Procedia Computer Science, 244, 342–352. https://doi.org/10.1016/j.procs.2024.10.208 DOI: https://doi.org/10.1016/j.procs.2024.10.208
Elgobshawi, A. E. (2024). Conceptualization of Morphological Roots in Arabic and English: A Contrastive Analysis. World Journal of English Language, 14(5), 436–443. https://doi.org/10.5430/wjel.v14n5p436 DOI: https://doi.org/10.5430/wjel.v14n5p436
Elkhbir, N. (2024). Information extraction for arabic and its dialects [Thesis, UNIVERSITÉ SORBONNE PARIS NORD]. https://doi.org/10.70675/702bb6e7z1efcz4563z9539zb485ffc11160 DOI: https://doi.org/10.70675/702bb6e7z1efcz4563z9539zb485ffc11160
Elnaka, A., Nael, O., Afifi, H., & Sharaf, N. (2021). AraScore: Investigating Response-Based Arabic Short Answer Scoring. Procedia CIRP, 189, 282–291. https://doi.org/10.1016/j.procs.2021.05.091 DOI: https://doi.org/10.1016/j.procs.2021.05.091
ElSabagh, A. A., Azab, S. S., & Hefny, H. A. (2025). A comprehensive survey on Arabic text augmentation: approaches, challenges, and applications. In Neural Computing and Applications. Springer Science and Business Media Deutschland GmbH. https://doi.org/10.1007/s00521-025-11020-z DOI: https://doi.org/10.1007/s00521-025-11020-z
Feng, S. Y., Gangal, V., Wei, J., Chandar, S., Vosoughi, S., Mitamura, T., & Hovy, E. (2021). A Survey of Data Augmentation Approaches for NLP. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 968–988. https://doi.org/10.18653/v1/2021.findings-acl.84 DOI: https://doi.org/10.18653/v1/2021.findings-acl.84
Gagliardi, I., & Artese, M. T. (2023). Ensemble-Based Short Text Similarity: An Easy Approach for Multilingual Datasets Using Transformers and WordNet in Real-World Scenarios. Big Data and Cognitive Computing, 7(4). https://doi.org/10.3390/bdcc7040158 DOI: https://doi.org/10.3390/bdcc7040158
Ghazoui, B., Bazi, I. El, Essadik, I., Benali, B. A., & Moussa, H. (2026). Robust Arabic tweet NER via label-aware data augmentation and AraBERTv2. Bulletin of Electrical Engineering and Informatics, 15(1), 740–754. https://doi.org/10.11591/eei.v15i1.10462 DOI: https://doi.org/10.11591/eei.v15i1.10462
Habbat, N., Nouri, H., Anoun, H., & Hassouni, L. (2023). Using AraGPT and ensemble deep learning model for sentiment analysis on Arabic imbalanced dataset. ITM Web of Conferences, 52, 02008. https://doi.org/10.1051/itmconf/20235202008 DOI: https://doi.org/10.1051/itmconf/20235202008
Habberrih, A., & Abuzaraida, M. A. (2024). A review of the available Arabic dialects datasets for Sentiment Analysis. Journal of Sustainable Research in Applied Sciences, 30(2), 30–37. DOI: https://doi.org/10.36602/jsras.2024.1.2.1
Hossain, M. M., Hossain, M. S., Safran, M., Alfarhood, S., Alfarhood, M., & Mridha, M. F. (2024). A Hybrid Attention-Based Transformer Model for Arabic News Classification Using Text Embedding and Deep Learning. IEEE Access, 12, 198046–198066. https://doi.org/10.1109/ACCESS.2024.3522061 DOI: https://doi.org/10.1109/ACCESS.2024.3522061
Hsu, T.-W., Chen, C.-C., Huang, H.-H., & Chen, H.-H. (2021). Semantics-Preserved Data Augmentation for Aspect-Based Sentiment Analysis. Conference on Empirical Methods in Natural Language Processing, 4417–4422. DOI: https://doi.org/10.18653/v1/2021.emnlp-main.362
Lee, J.-M., & Ha, T.-B. (2023). Unsupervised Text Embedding Space Generation Using Generative Adversarial Networks for Text Synthesis. https://doi.org/10.3384/nejlt.2000-1533.2023.4855 DOI: https://doi.org/10.3384/nejlt.2000-1533.2023.4855
Li, B., Hou, Y., & Che, W. (2022). Data augmentation approaches in natural language processing: A survey. AI Open, 3, 71–90. https://doi.org/10.1016/j.aiopen.2022.03.001 DOI: https://doi.org/10.1016/j.aiopen.2022.03.001
Mohamed, E. A., Ismail, W. N., Ibrahim, O. A. S., & Younis, E. M. G. (2024). A two-stage framework for Arabic social media text misinformation detection combining data augmentation and AraBERT. Social Network Analysis and Mining, 14(1). https://doi.org/10.1007/s13278-024-01201-4 DOI: https://doi.org/10.1007/s13278-024-01201-4
Mumuni, A., & Mumuni, F. (2022). Data augmentation: A comprehensive survey of modern approaches. In Array (Vol. 16). Elsevier B.V. https://doi.org/10.1016/j.array.2022.100258 DOI: https://doi.org/10.1016/j.array.2022.100258
Nacar, O., Sibaee, S., Ahmed, S., Alharbi, A. I., Ghouti, L., & Koubaa, A. (2024). ASOS at OSACT6 Shared Task: Investigation of Data Augmentation in Arabic Dialect-MSA Translation. ASOS at OSACT6 Shared Task: Investigation of Data Augmentation in Arabic Dialect-MSA Translation, 104–111. DOI: https://doi.org/10.63317/34u3y58xvz4e
Refai, D., Abu-Soud, S., & Abdel-Rahman, M. J. (2023a). Data Augmentation Using Transformers and Similarity Measures for Improving Arabic Text Classification. IEEE Access, 11, 132516–132531. https://doi.org/10.1109/ACCESS.2023.3336311
Refai, D., Abu-Soud, S., & Abdel-Rahman, M. J. (2023b). Data Augmentation Using Transformers and Similarity Measures for Improving Arabic Text Classification. IEEE Access, 11, 132516–132531. https://doi.org/10.1109/ACCESS.2023.3336311 DOI: https://doi.org/10.1109/ACCESS.2023.3336311
Sabty, C., Omar, I., Wasfalla, F., Islam, M., & Abdennadher, S. (2021). Data Augmentation Techniques on Arabic Data for Named Entity Recognition. Procedia Computer Science, 189, 292–299. https://doi.org/10.1016/j.procs.2021.05.092 DOI: https://doi.org/10.1016/j.procs.2021.05.092
Talafha, B., Fadel, A., Al-Ayyoub, M., Jararweh, Y., Al-Smadi, M., & Juola, P. (2019). Team JUST at the MADAR Shared Task on Arabic Fine-Grained Dialect Identification. Proceedings of the Fourth Arabic Natural Language Processing, 285–289. https://doi.org/10.18653/v1/W19-4638 DOI: https://doi.org/10.18653/v1/W19-4638
Worth, P. J. (2023). Word Embeddings and Semantic Spaces in Natural Language Processing. International Journal of Intelligence Science, 13(01), 1–21. https://doi.org/10.4236/ijis.2023.131001 DOI: https://doi.org/10.4236/ijis.2023.131001
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 علي مفتاح بن عمران ، أشرف علي ناصف، أيمن مختار ارميص، سالم حسين المدهون، معمر مصباح عوينات

This work is licensed under a Creative Commons Attribution 4.0 International License.



