Predicting Emerging Cyber Threats Using Natural Language Processing (NLP)

Authors

  • Shayma Jawad Ministry of Education/Third Rusafa Directorate iraq
  • Hiba Hameed majeed Ministry of Education/Third Rusafa Directorate iraq

DOI:

https://doi.org/10.71229/acr6zw19

Keywords:

Cyber Threat Intelligence(CTI), , Natural Language,, Processing (NLP), , Bi-LSTM,, Random Forest, , Malware Classification,

Abstract

The rapid evolution of cyber-threats, including Advanced Persistent Threats (APTs) and dynamic malware strains, has made signature-based defense techniques obsolete. Proactive Cyber Threat Intelligence (CTI) and autonomous host-based detection have emerged as critical paradigms to circumvent modern threats before any system is exploited. This paper proposes a full-fledged, multi-level framework that approaches security telemetry and unstructured textual logs like natural language constructs. We evaluate our system on a heterogeneous dataset (10,000 samples containing 8 different threat domains: adware, botnet, phishing, ransomware, rootkit, spyware, Trojan, and worm) and introduce a dual-engine architecture: an Early-Stage Structural Identifier with static metadata feature engineering, using together an ensemble bagging tree classifier as predictor, and a Linguistic Context Predictor supported by Bidirectional Long Short-Term Memory (Bi-LSTM) networks and customized Transformers. In an empirical study, this was revealed to perform best when the structural metadata, which has high cross-family structural entropy, produced an accuracy baseline of 23.15%, while the linguistic engine that processes text-mined security intents reached an accuracy rate of up to 100.00% with a false positive rate of 0.00%. This leads to a scalable approach toward modern enterprise ecosystems that support resilient, real-time autonomous threat mitigation.

References

[1] Ahmadou, F., Ghaffarzadegan, S., Nour, B., Pourzandi, M., Debbabi, M., & Assi, C. (2025, July). Automated attack testflow extraction from cyber threat report using BERT for contextual analysis. arXiv. https://arxiv.org/abs/2507.07244 DOI: https://doi.org/10.1109/NoF66640.2025.11223332

[2] Almahmoud, Z., Yoo, P. D., Alhussein, O., Farhat, I., & Damiani, E. (2023, May). A holistic and proactive approach to forecasting cyber threats. Scientific Reports, 13, Article 8049. DOI: https://doi.org/10.1038/s41598-023-35198-1

[3] Bayer, M., Kuehn, P., Shanehsaz, R., & Reuter, C. (2022, December). CYSECBERT: A domain-adapted language model for the cybersecurity domain. arXiv. https://arxiv.org/abs/2212.02974

[4] Cotroneo, D., Natella, R., & Orbinato, V. (2025, May). Elevating cyber threat intelligence against disinformation campaigns with LLM-based concept extraction and the FakeCTI dataset. arXiv. https://arxiv.org/abs/2505.03345 DOI: https://doi.org/10.2139/ssrn.5202513

[5] Ferrag, M. A., Ndhlovu, M., Tihanyi, N., Cordeiro, L. C., Debbah, M., & Lestable, T. (2023, June). Revolutionizing cyber threat detection with large language models. arXiv. https://arxiv.org/abs/2306.14263

[6] Gong, C., Li, Z., & Li, X. (2025, July). Information security based on LLM approaches: A review. arXiv. https://arxiv.org/abs/2507.18215

[7] Ismail, W. S. (2024, February). Threat detection and response using AI and NLP in cybersecurity. Journal of Internet Services and Information Security (JISIS), 14(1), 195–205. DOI: https://doi.org/10.58346/JISIS.2024.I1.013

[8] Jamadi, Z., & Aghdam, A. G. (2023, June). Early malware detection and next-action prediction. arXiv. https://arxiv.org/abs/2306.06255

[9] Kamdem, I. G. K., & Nkenlifack, M. (2024, April). Cyber deception using NLP. Journal of Information Security, 15(2), 279–297. DOI: https://doi.org/10.4236/jis.2024.152016

[10] Kheddar, H. (2024). Transformers and large language models for efficient intrusion detection systems: A comprehensive survey (Preprint). University of Medea. DOI: https://doi.org/10.1016/j.inffus.2025.103347

[11] Melon, M. M. H., et al. (2025, September). Phishing detection in the age of NLP: Leveraging deep and machine learning for enhanced accuracy. Journal of Advanced Research in Applied Sciences and Engineering Technology, 55(2), 264–273. DOI: https://doi.org/10.37934/araset.55.2.264273

[12] Naseer, I. (2024, January). Machine learning applications in cyber threat intelligence: A comprehensive review. Asian Bulletin of Big Data Management, 3(2), 190–200. DOI: https://doi.org/10.62019/abbdm.v3i2.85

[13] Pal, K. K., Kuznia, K. C., Kashihara, K., Jagtap, S., Anantheswaran, U., & Baral, C. (2023, February). Exploring the limits of transfer learning with unified model in the cybersecurity domain. arXiv. https://arxiv.org/abs/2302.10346

[14] Rhode, M., Burnap, P., & Jones, K. (2018). Early-stage malware prediction using recurrent neural networks. Computers & Security, 77, 578–594. DOI: https://doi.org/10.1016/j.cose.2018.05.010

[15] Shafee, S., Bessani, A., & Ferreira, P. M. (2025, July). False alarms, real damage: Adversarial attacks using LLM-based models on text-based cyber threat intelligence systems. arXiv. https://arxiv.org/abs/2507.06252 DOI: https://doi.org/10.1016/j.future.2026.108603

[16] Sorokoletova, O., Antonioni, E., & Colò, G. (2023). Towards a scalable AI-driven framework for data-independent cyber threat intelligence information extraction. In IEEE Proceedings. CY4GATE S.p.A. DOI: https://doi.org/10.1109/FLLM63129.2024.10852465

[17] Stehr, M.-O., & Kim, M. (2023, August). Vulnerability clustering and other machine learning applications of semantic vulnerability embeddings. arXiv. https://arxiv.org/abs/2310.05935

[18] Sufi, F. (2023, March). A new social media-driven cyber threat intelligence. Electronics, 12(5), Article 1242. DOI: https://doi.org/10.3390/electronics12051242

[19] Sworna, Z. T., Mousavi, Z., & Babar, M. A. (2022, November). NLP methods in host-based intrusion detection systems: A systematic review and future directions. Journal of Network and Computer Applications, 207, Article 103444. DOI: https://doi.org/10.1016/j.jnca.2023.103761

[20] Zimba, A., & Phiri, K. O. (2025, October). An enhanced machine learning with NLP modelling technique for smishing attacks detection in low-resourced languages. Research Square. DOI: https://doi.org/10.2139/ssrn.5195337

Downloads

Published

2026-09-29

Issue

Section

Original Articles

How to Cite

Predicting Emerging Cyber Threats Using Natural Language Processing (NLP). (2026). Al-Noor Journal of Engineering Management and Computer Science, 3(1), 63-67. https://doi.org/10.71229/acr6zw19

Similar Articles

81-90 of 100

You may also start an advanced similarity search for this article.