2 matches found
relay-secreport-bench
secreport-bench — the Security-Report Benchmark An evaluation...
5.9AI score
SaveExploits0
The Steganographic Potentials of Language Models
The potential for large language models LLMs to hide messages within plain text steganography poses a challenge to detection and thwarting of unaligned AI agents, and undermines faithfulness of LLMs reasoning. We explore the steganographic capabilities of LLMs fine-tuned via reinforcement learnin...
6.7AI score
SaveExploits0
20