Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
•added 2026/07/01 12:00 a.m.•16 views

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety

Large language models LLMs can be induced to produce harmful content through multi turn strategies in which no single user message appears clearly unsafe. Existing runtime safeguards commonly evaluate prompts or responses as isolated messages, which limits their ability to recover ac-cumulated...

6AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/05/22 12:00 a.m.•18 views

MTSA: Multi-Turn Safety Alignment for LLMs through Multi-Round Red-Teaming

Whitepaper called MTSA: Multi-Turn Safety Alignment For LLMs Through Multi-Round Red-Teaming...

7AI score
SaveExploits0
Rows per page
Query Builder