1 matches found
Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection
Large language models LLMs are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation through prompt injection and jailbreak attacks. Classifier-based guardrails, such as Prompt Guard 2, are widely used as a first line of defense against suc...
5.8AI score
SaveExploits0
20