Real-Time Detection and Mitigation of Prompt Injection Attacks in LLM-Integrated Enterprise Systems

Authors

  • Fatimah Ghazi Suwaidan Department of Cybersecurity, College of Computer Science and Information Technology, University of Al-Qadisiyah, Al Diwaniyah, Iraq,

DOI:

https://doi.org/10.71229/5cn8a439

Keywords:

Prompt Injection, LLM Security, Real-Time Threat Detection, Enterprise AI Systems, Adaptive Guardrails, Explainable Security, Retrieval-Augmented Generation Security, Agentic AI Safety

Abstract

Large language models (LLMs) embedded in enterprise workflows cannot structurally distinguish legitimate instructions from adversarial ones in the same token stream, making prompt injection OWASP's top LLM risk for two consecutive editions a persistent threat across direct and indirect vectors. This paper presents PromptShield-RT, a layered, real-time, model-agnostic framework combining input normalization and provenance tagging, lexical-heuristic pattern matching, a statistical classifier, structural anomaly features, and calibrated risk fusion, with policy-driven mitigation (allow/sanitize/quarantine/block) and an explainable, adaptive-feedback mechanism for SOC workflows. We construct an original evaluation corpus, SynPI-Bench (n = 450, six categories), and a template-disjoint held-out generalization set (n = 31) with novel phrasings, obfuscation encodings, and adversarial hard-negative benign text. Using template-grouped 5-fold cross-validation, the fused pipeline achieves 92.4% accuracy (F1 = 0.930, AUC = 0.990), outperforming heuristic-only (57.0%) and naive-averaged (59.2%) baselines, while a lexical classifier reaches 85.9% with lower precision. We report a pronounced generalization gap on the held-out set (48.4% accuracy, 90% false-positive rate on hard negatives), quantifying a known limitation of surface-lexical defenses. The pipeline achieves sub-millisecond P95 latency (0.266 ms), within typical 50 ms enterprise SLAs. We situate PromptShield-RT relative to structural, architectural, and guardrail-product defenses, arguing for layered, defense-in-depth architectures, with reproducible code provided.

References

[1] F. Perez and I. Ribeiro, "Ignore previous prompt: Attack techniques for language models," in NeurIPS ML Safety Workshop, 2022.

[2] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, "Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection," in Proc. 16th ACM Workshop on Artificial Intelligence and Security (AISec), 2023.

[3] S. Chen, J. Piet, C. Sitawarin, and D. Wagner, "StruQ: Defending against prompt injection with structured queries," in Proc. 34th USENIX Security Symposium, Seattle, WA, 2025, pp. 2383–2400.

[4] S. Chen, A. Zharmagambetov, S. Mahloujifar, K. Chaudhuri, D. Wagner, and C. Guo, "SecAlign: Defending against prompt injection with preference optimization," in Proc. 2025 ACM SIGSAC Conf. on Computer and Communications Security, 2025, pp. 2833–2847.

[5] E. Wallace, K. Xiao, R. Leike, L. Weng, J. Heidecke, and A. Beutel, "The instruction hierarchy: Training LLMs to prioritize privileged instructions," arXiv:2404.13208, 2024.

[6] T. Wu, S. Zhang, K. Song, S. Xu, S. Zhao, R. Agrawal, S. R. Indurthi, C. Xiang, P. Mittal, and W. Zhou, "Instructional segment embedding: Improving LLM safety with instruction hierarchy," arXiv:2410.09102, 2024.

[7] J. Piet, M. Alrashed, C. Sitawarin, S. Chen, Z. Wei, E. Sun, B. Alomair, and D. Wagner, "Jatmo: Prompt injection defense by task-specific finetuning," in Proc. European Symp. on Research in Computer Security (ESORICS), 2024, pp. 105–124.

[8] M. Sharma, M. Tong, J. Mu, J. Wei, J. Kruthoff, S. Goodfriend, E. Ong, A. Peng, R. Agarwal, C. Anil, et al., "Constitutional classifiers: Defending against universal jailbreaks across thousands of hours of red teaming," arXiv:2501.18837, 2025.

[9] E. Debenedetti, J. Zhang, M. Balunovic, L. Beurer-Kellner, M. Fischer, and F. Tramèr, "Defeating prompt injections by design (CaMeL)," arXiv:2503.18813, 2025.

[10] S. Willison, "The dual LLM pattern for building AI assistants that can resist prompt injection," simonwillison.net, 2023.

[11] M. Costa, B. Köpf, A. Kolluri, A. Paverd, M. Russinovich, A. Salem, S. Tople, L. Wutschitz, and S. Zanella-Béguelin, "Securing AI agents with information-flow control," arXiv:2505.23643, 2025.

[12] "MELON: Provable defense against indirect prompt injection attacks in AI agents," arXiv:2502.05174, 2025.

[13] A. Robey, E. Wong, H. Hassani, and G. J. Pappas, "SmoothLLM: Defending large language models against jailbreaking attacks," arXiv:2310.03684, 2023.

[14] Protect AI, "deberta-v3-base-prompt-injection-v2: A prompt-injection detection classifier," Hugging Face model card, 2024.

[15] T. Rebedea, R. Dinu, M. N. Sreedhar, C. Parisien, and J. Cohen, "NeMo Guardrails: A toolkit for controllable and safe LLM applications with programmable rails," in Proc. 2023 Conf. on Empirical Methods in Natural Language Processing: System Demonstrations, 2023.

[16] M. A. Rahman, H. Shahriar, G. Francia, F. Wu, A. Cuzzocrea, M. Rahman, M. J. Faruk, and S. Ahamed, "Fine-tuned large language models (LLMs): Improved prompt injection attacks detection," 2025.

[17] A. Razavi, M. Soltangheis, N. Arabzadeh, S. Salamat, M. Zihayat, and E. Bagheri, "Benchmarking prompt sensitivity in large language models," in Proc. European Conf. on Information Retrieval (ECIR), Springer, 2025, pp. 303–313.

[18] S. Toyer, O. Watkins, E. A. Mendes, J. Svegliato, L. Bailey, T. Wang, I. Ong, K. Elmaaroufi, P. Abbeel, T. Darrell, A. Ritter, and S. Russell, "Tensor Trust: Interpretable prompt injection attacks from an online game," in Proc. Int. Conf. on Learning Representations (ICLR), 2024.

[19] E. Debenedetti, J. Zhang, M. Balunovic, L. Beurer-Kellner, M. Fischer, and F. Tramèr, "AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents," in Proc. NeurIPS 2024 Datasets and Benchmarks Track, 2024.

[20] OWASP Foundation, "OWASP Top 10 for LLM Applications 2025," 2025. Available: https://owasp.org/www-project-top-10-for-large-language-model-applications/assets/PDF/OWASP-Top-10-for-LLMs-v2025.pdf

[21] National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)," 2024.

[22] Cisco, "Prompt injection is the new SQL injection, and guardrails aren't enough," Cisco Blogs, Mar. 2026.

[23] "Are AI-assisted development tools immune to prompt injection?," arXiv:2603.21642, 2026.

[24] T. Geng, Z. Xu, Y. Qu, and W. E. Wong, "Prompt injection attacks on large language models: A survey of attack methods, root causes, and defense strategies," Computers, Materials & Continua, vol. 87, no. 1, p. 4, 2026, doi: 10.32604/cmc.2025.074081.

[25] "Hijacking the prompt: A survey of prompt injection attacks, detection, and defense in large language models," Cybersecurity Undergraduate Research Showcase, Old Dominion University Digital Commons, 2026.

[26] "Prompt injection attacks in large language models and AI agent systems: A comprehensive review of vulnerabilities, attack vectors, and defense mechanisms," Information (MDPI), vol. 17, no. 1, art. 54, 2026.

fig 3

Downloads

Published

2026-08-09

Issue

Section

Original Articles

How to Cite

Real-Time Detection and Mitigation of Prompt Injection Attacks in LLM-Integrated Enterprise Systems. (2026). Al-Noor Journal of Engineering Management and Computer Science, 2(3), 99-113. https://doi.org/10.71229/5cn8a439