Data Augmentation Techniques for Arabic Short Text Processing
Comparative Insights and Future Directions
DOI:
https://doi.org/10.59743/jaf.v10i1.960Keywords:
Arabic NLP, Data Augmentation, Short Text Processing, Dialectal Variation, Low-Resource Languages, Comparative ReviewAbstract
Arabic short texts pose challenges due to morphological complexity, dialectal diversity, and data scarcity. Data augmentation offers a solution, yet no review systematically compares techniques across tasks. We examine six strategies across five tasks using 22 studies (2019–2026). Results show no universal method works: synonym replacement boosts topic classification but harms NER; back-translation fails under dialectal shift; generative models improve sentiment analysis when filtered; feature-space resampling inflates metrics artificially. Three persistent issues cut across tasks: the diversity–fidelity trade-off, dialectal contamination, and computational cost. We propose a focused agenda emphasizing contrastive sampling, bias-aware generation, multimodal integration, and lightweight augmentation for low-resource dialects. This review provides actionable guidance for task-aware Arabic NLP frameworks.
Downloads
References
Abdhood, S. F., Omar, N., & Tiun, S. (2025). A Novel Data Augmentation Framework for Arabic Multi-Label Text Classification Using AraBART, AraGPT2, and Borderline-SMOTE. IEEE Access, 13, 169769–169778. https://doi.org/10.1109/ACCESS.2025.3609462
Alabdullah, A., Han, L., & Lin, C. (2025). Advancing Dialectal Arabic to Modern Standard Arabic Machine Translation. https://doi.org/10.21203/rs.3.rs-7510599/v1
Alayba, A. M. (2025). Arabic Natural Language Processing (NLP): A Comprehensive Review of Challenges, Techniques, and Emerging Trends. In Computers (Vol. 14, Number 11). Multidisciplinary Digital Publishing Institute (MDPI). https://doi.org/10.3390/computers14110497
Aljohani, E. (2024). Enhancing Arabic Fake News Detection: Evaluating Data Balancing Techniques Across Multiple Machine Learning Models. Engineering, Technology and Applied Science Research, 14(4), 15947–15956. https://doi.org/10.48084/etasr.8019
Al-Khazaleh, M. J., Alian, M., & Jaradat, M. A. (2024). Sentiment analysis of imbalanced Arabic data using sampling techniques and classification algorithms. Bulletin of Electrical Engineering and Informatics, 13(1), 607–618. https://doi.org/10.11591/eei.v13i1.5886
Almutairi, S., & Alotaibi, F. (2023). A Comparative Analysis for Arabic Sentiment Analysis Models In E-Marketing Using Deep Learning Techniques. In Journal of Engineering and Applied Sciences (Vol. 10, Number 1).
Alotaibi, A., & Nadeem, F. (2024). Leveraging Social Media and Deep Learning for Sentiment Analysis for Smart Governance: A Case Study of Public Reactions to Educational Reforms in Saudi Arabia. Computers, 13(11). https://doi.org/10.3390/computers13110280
Alrashidi, B., Jamal, A., & Alkhathlan, A. (2023). Abusive Content Detection in Arabic Tweets Using Multi-Task Learning and Transformer-Based Models. Applied Sciences (Switzerland), 13(10). https://doi.org/10.3390/app13105825
Badri, N., Kboubi, F., & Habacha Chaibi, A. (2024). Abusive and Hate speech Classification in Arabic Text Using Pre-trained Language Models and Data Augmentation. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(11). https://doi.org/10.1145/3679049
Bayer, M., Kaufhold, M. A., & Reuter, C. (2022). A Survey on Data Augmentation for Text Classification. ACM Computing Surveys, 55(7). https://doi.org/10.1145/3544558
Chen, S., Zhang, Y., & Yang, Q. (2024). Multi-Task Learning in Natural Language Processing: An Overview. ACM Computing Surveys, 56(12). https://doi.org/10.1145/3663363
Cooper, R., Kliesner, K. W., & Zenker, S. (2024). Contrastive Meta-Learner for Automatic Text Labeling and Semantic Textual Similarity. IEEE Access, 12, 166792–166799. https://doi.org/10.1109/ACCESS.2024.3424401
Dahou, A. H., Cheragui, M. A., Abdedaiem, A., & Mathiak, B. (2024). Enhancing Model Performance through Translation-based Data Augmentation in the context of Fake News Detection. Procedia Computer Science, 244, 342–352. https://doi.org/10.1016/j.procs.2024.10.208
Elgobshawi, A. E. (2024). Conceptualization of Morphological Roots in Arabic and English: A Contrastive Analysis. World Journal of English Language, 14(5), 436–443. https://doi.org/10.5430/wjel.v14n5p436
Elkhbir, N. (2024). Information extraction for arabic and its dialects [Thesis, UNIVERSITÉ SORBONNE PARIS NORD]. https://doi.org/10.70675/702bb6e7z1efcz4563z9539zb485ffc11160
Elnaka, A., Nael, O., Afifi, H., & Sharaf, N. (2021). AraScore: Investigating Response-Based Arabic Short Answer Scoring. Procedia CIRP, 189, 282–291. https://doi.org/10.1016/j.procs.2021.05.091
ElSabagh, A. A., Azab, S. S., & Hefny, H. A. (2025). A comprehensive survey on Arabic text augmentation: approaches, challenges, and applications. In Neural Computing and Applications. Springer Science and Business Media Deutschland GmbH. https://doi.org/10.1007/s00521-025-11020-z
Feng, S. Y., Gangal, V., Wei, J., Chandar, S., Vosoughi, S., Mitamura, T., & Hovy, E. (2021). A Survey of Data Augmentation Approaches for NLP. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 968–988. https://doi.org/10.18653/v1/2021.findings-acl.84
Gagliardi, I., & Artese, M. T. (2023). Ensemble-Based Short Text Similarity: An Easy Approach for Multilingual Datasets Using Transformers and WordNet in Real-World Scenarios. Big Data and Cognitive Computing, 7(4). https://doi.org/10.3390/bdcc7040158
Ghazoui, B., Bazi, I. El, Essadik, I., Benali, B. A., & Moussa, H. (2026). Robust Arabic tweet NER via label-aware data augmentation and AraBERTv2. Bulletin of Electrical Engineering and Informatics, 15(1), 740–754. https://doi.org/10.11591/eei.v15i1.10462
Habbat, N., Nouri, H., Anoun, H., & Hassouni, L. (2023). Using AraGPT and ensemble deep learning model for sentiment analysis on Arabic imbalanced dataset. ITM Web of Conferences, 52, 02008. https://doi.org/10.1051/itmconf/20235202008
Habberrih, A., & Abuzaraida, M. A. (2024). A review of the available Arabic dialects datasets for Sentiment Analysis. Journal of Sustainable Research in Applied Sciences, 30(2), 30–37.
Hossain, M. M., Hossain, M. S., Safran, M., Alfarhood, S., Alfarhood, M., & Mridha, M. F. (2024). A Hybrid Attention-Based Transformer Model for Arabic News Classification Using Text Embedding and Deep Learning. IEEE Access, 12, 198046–198066. https://doi.org/10.1109/ACCESS.2024.3522061
Hsu, T.-W., Chen, C.-C., Huang, H.-H., & Chen, H.-H. (2021). Semantics-Preserved Data Augmentation for Aspect-Based Sentiment Analysis. Conference on Empirical Methods in Natural Language Processing, 4417–4422.
Lee, J.-M., & Ha, T.-B. (2023). Unsupervised Text Embedding Space Generation Using Generative Adversarial Networks for Text Synthesis. https://doi.org/10.3384/nejlt.2000-1533.2023.4855
Li, B., Hou, Y., & Che, W. (2022). Data augmentation approaches in natural language processing: A survey. AI Open, 3, 71–90. https://doi.org/10.1016/j.aiopen.2022.03.001
Mohamed, E. A., Ismail, W. N., Ibrahim, O. A. S., & Younis, E. M. G. (2024). A two-stage framework for Arabic social media text misinformation detection combining data augmentation and AraBERT. Social Network Analysis and Mining, 14(1). https://doi.org/10.1007/s13278-024-01201-4
Mumuni, A., & Mumuni, F. (2022). Data augmentation: A comprehensive survey of modern approaches. In Array (Vol. 16). Elsevier B.V. https://doi.org/10.1016/j.array.2022.100258
Nacar, O., Sibaee, S., Ahmed, S., Alharbi, A. I., Ghouti, L., & Koubaa, A. (2024). ASOS at OSACT6 Shared Task: Investigation of Data Augmentation in Arabic Dialect-MSA Translation. ASOS at OSACT6 Shared Task: Investigation of Data Augmentation in Arabic Dialect-MSA Translation, 104–111.
Refai, D., Abu-Soud, S., & Abdel-Rahman, M. J. (2023a). Data Augmentation Using Transformers and Similarity Measures for Improving Arabic Text Classification. IEEE Access, 11, 132516–132531. https://doi.org/10.1109/ACCESS.2023.3336311
Refai, D., Abu-Soud, S., & Abdel-Rahman, M. J. (2023b). Data Augmentation Using Transformers and Similarity Measures for Improving Arabic Text Classification. IEEE Access, 11, 132516–132531. https://doi.org/10.1109/ACCESS.2023.3336311
Sabty, C., Omar, I., Wasfalla, F., Islam, M., & Abdennadher, S. (2021). Data Augmentation Techniques on Arabic Data for Named Entity Recognition. Procedia Computer Science, 189, 292–299. https://doi.org/10.1016/j.procs.2021.05.092
Talafha, B., Fadel, A., Al-Ayyoub, M., Jararweh, Y., Al-Smadi, M., & Juola, P. (2019). Team JUST at the MADAR Shared Task on Arabic Fine-Grained Dialect Identification. Proceedings of the Fourth Arabic Natural Language Processing, 285–289. https://doi.org/10.18653/v1/W19-4638
Worth, P. J. (2023). Word Embeddings and Semantic Spaces in Natural Language Processing. International Journal of Intelligence Science, 13(01), 1–21. https://doi.org/10.4236/ijis.2023.131001
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 علي مفتاح بن عمران ، أشرف علي ناصف، أيمن مختار ارميص، سالم حسين المدهون، معمر مصباح عوينات

This work is licensed under a Creative Commons Attribution 4.0 International License.



