2 matches found
Quant Fever, Reasoning Blackholes, Schrodinger'S Compliance, and More: Probing GPT-OSS-20B
OpenAI's GPT-OSS family provides open-weight language models with explicit chain-of-thought CoT reasoning and a Harmony prompt format. We summarize an extensive security evaluation of GPT-OSS-20B that probes the model's behavior under different adversarial conditions. Using the Jailbreak Oracle J...
6.8AI score
SaveExploits0
LLM Jailbreak Oracle
As large language models LLMs become increasingly deployed in safety-critical applications, the lack of systematic methods to assess their vulnerability to jailbreak attacks presents a critical security gap. We introduce the jailbreak oracle problem: given a model, prompt, and decoding strategy,...
7.4AI score
SaveExploits0
20