Lucene search
+L

2 matches found

Kitploit
Kitploit
•added 2026/10/06 3:58 a.m.•1 views

benign-instruction-bench

Passing the Test You Trained On Re-evaluating prompt-injection detectors where LLM agents actually use them: on the tool outputs an agent reads. Teams pick injection detectors by their benchmark scores. We check whether those scores predict behavior inside an agent, and find that they mostly...

SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/11 12:00 a.m.•11 views

Effective Red-Teaming of Policy-Adherent Agents

Task-oriented LLM-based agents are increasingly used in domains with strict policies, such as refund eligibility or cancellation rules. The challenge lies in ensuring that the agent consistently adheres to these rules and policies, appropriately refusing any request that would violate them, while...

6.9AI score
SaveExploits0
Rows per page
Query Builder