Selective Arabic Cyberbullying Detection with Multi-View Modeling, Risk Triage, and Human Deferral

Authors

  • Mohammed A.S Al-Hitawi Department of Artificial Intelligence, College of Information Technology, University of Fallujah, 31002, Fallujah, Anbar, Iraq https://orcid.org/0009-0009-7905-0978

DOI:

https://doi.org/10.71229/er3kch93

Keywords:

Arabic cyberbullying, Selective classification, Human-in-the-loop moderation, Risk triage, Arabic language understanding

Abstract

The identification of Arabic cyberbullying is generally perceived as a closed-set categorization challenge, but moderation tools need to address cases that are ambiguous, cases involving domain change in situations when review priorities are not equal. Current research introduces a selective moderation approach which can produce only either a decision about cyberbullying or refer the case to human reviewers with an optional high-priority triage flag of abusive post. As a main experiment employs a public Arabic cyberbullying corpus that was transformed into 13230 logical records and subsequently reduced to 12961 unique normalized texts after deleting 269 duplicate records (6739 cyberbullying and 6222 non-cyberbullying). Word unigram/bigram and character n-grams 3–5 were organized into TF-IDF representation, concatenated, and classified using the class-balanced LinearSVC. The results of five-fold stratified cross-validation show F1=0.9914±0.0013, macro-F1=0.9911±0.0014 and ROC-AUC=0.9996±0.0001 (mean±SD). Eventually, there was an uncertainty reject option that forwarded all cases where there was no high confidence to human agents. The deferral of 5% of the cases leads to capturing 85.7% of the errors made in non-deferral and improves the retained case’s F1 to 0.9987; a deferral of 10% of the cases captures 94.6% of the errors made and results in a retained-case F1 of 0.9995. Additionally, two audits are used to avoid inaccurate interpretations of the within-corpus score. The first involves training the cyber-bullying corpus and testing on 4,000-comment offensive language dataset of MPOLD which reduces the F1 score to 0.3882 (ROC-AUC = 0.7318). The second audit uses a different subtype proxy of hate speech from the MPOLD text allowing to obtain F1= 0.5771 and ROC-AUC = 0.7996 while addressing offensive items. Thus, it provides evidence of a necessity of human prioritization instead of autonomous risk making decisions.

References

[1] M. Khairy, T. M. Mahmoud, and T. Abd-El-Hafeez, “Automatic detection of cyberbullying and abusive language in Arabic content on social networks: A survey,” Procedia Computer Science, vol. 189, pp. 156–166, 2021, doi: 10.1016/j.procs.2021.05.080.

[2] H. Allwaibed, M. Anbar, S. Manickam, and A. Bintang, “Cyberbullying detection approaches for Arabic texts: A systematic literature review,” Frontiers in Artificial Intelligence, vol. 8, art. 1666349, 2025, doi: 10.3389/frai.2025.1666349.

[3] F. Shannag, B. H. Hammo, and H. Faris, “The design, construction and evaluation of annotated Arabic cyberbullying corpus,” Education and Information Technologies, vol. 27, pp. 10977–11023, 2022, doi: 10.1007/s10639-022-11056-x.

[4] A. M. Alduailaj and A. Belghith, “Detecting Arabic Cyberbullying Tweets Using Machine Learning,” Machine Learning and Knowledge Extraction, vol. 5, no. 1, pp. 29–42, 2023, doi: 10.3390/make5010003.

[5] M. Alzaqebah, G. M. Jaradat, D. Nassan, R. Alnasser, M. K. Alsmadi, I. Almarashdeh, and S. Alkhushayni, “Cyberbullying detection framework for short and imbalanced Arabic datasets,” Journal of King Saud University - Computer and Information Sciences, vol. 35, no. 8, art. 101652, 2023, doi: 10.1016/j.jksuci.2023.101652.

[6] M. Khairy, T. M. Mahmoud, A. Omar, and T. Abd El-Hafeez, “Comparative performance of ensemble machine learning for Arabic cyberbullying and offensive language detection,” Language Resources and Evaluation, vol. 58, pp. 695–712, 2024, doi: 10.1007/s10579-023-09683-y.

[7] M. Azzeh, B. Alhijawi, A. Tabbaza, O. Alabboshi, N. Hamdan, D. Jaser, et al., “Arabic cyberbullying detection system using convolutional neural network and multi-head attention,” International Journal of Speech Technology, vol. 27, pp. 521–537, 2024, doi: 10.1007/s10772-024-10118-4.

[8] M. A. Mahdi, S. M. Fati, M. A. G. Hazber, S. Ahamad, and S. A. Saad, “Enhancing Arabic Cyberbullying Detection with End-to-End Transformer Model,” Computer Modeling in Engineering & Sciences, vol. 141, no. 2, pp. 1651–1671, 2024, doi: 10.32604/cmes.2024.052291.

[9] R. I. Hithnawi, A. Hamarsheh, and M. Maree, “AraBERT for Arabic cyberbullying detection in Facebook comments,” Journal of Cybersecurity, vol. 11, no. 1, art. tyaf030, 2025, doi: 10.1093/cybsec/tyaf030.

[10] A. M. Eissa, S. K. Guirguis, and M. M. Madbouly, “An optimized Arabic cyberbullying detection approach based on genetic algorithms,” Scientific Reports, vol. 15, art. 38479, 2025, doi: 10.1038/s41598-025-23586-8.

[11] W. Antoun, F. Baly, and H. Hajj, “AraBERT: Transformer-based Model for Arabic Language Understanding,” in Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, 2020, pp. 9–15.

[12] M. Abdul-Mageed, A. Elmadany, and E. M. B. Nagoudi, “ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, 2021, pp. 7088–7105.

[13] Y. Geifman and R. El-Yaniv, “SelectiveNet: A Deep Neural Network with an Integrated Reject Option,” in Proceedings of the 36th International Conference on Machine Learning, PMLR 97, 2019, pp. 2151–2159.

[14] S. A. Chowdhury, H. Mubarak, A. Abdelali, S.-g. Jung, B. J. Jansen, and J. Salminen, “A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection,” in Proceedings of the 12th Language Resources and Evaluation Conference, 2020, pp. 6203–6212.

[15] Al-Hitawi, M. A., & Gyöngyössy, N. M. (2026). Enhancing Transformer-Based Language Models for Hungarian Handwritten Text Recognition. F1000Research, 15, 181.

[16] Frhan, A. J., Al-Hitawi, M. A. S., & Abu-Alsaad, H. A. (2026). Supervised fine-tuning approach for medical question-answering using Qwen instruct. International Journal of Computers, 11, 57–62.

[17] الربيعي ن. م., “The Mechanisms of Obtaining Digital Evidence and using it as Proof in Cybercrimes”, Res. J. Legal Sci., vol. 5, no. 2, pp. 170–, Dec. 2024.

fig 2

Downloads

Published

2026-09-01

Issue

Section

Original Articles

How to Cite

Selective Arabic Cyberbullying Detection with Multi-View Modeling, Risk Triage, and Human Deferral. (2026). Al-Noor Journal of Engineering Management and Computer Science, 2(4), 25-36. https://doi.org/10.71229/er3kch93