3 matches found
Research on Models Engaging in Genie-Like Behavior
New paper: "Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training." Abstract: We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models RLMs, which we call self-jailbreaking. Specifically,...
Malicious code in rlms (npm)
--- -= Per source details. Do not edit below this line.=- Source: ghsa-malware 2335caf4ec81ea00b6da46ec70fa9da3bd1887c3d87dd1fc24a9466f45f98e0e Any computer that has this package installed or running should be considered fully compromised. All secrets and keys stored on that computer should be...
MAL-2022-5820 Malicious code in rlms (npm)
--- -= Per source details. Do not edit below this line.=- Source: ghsa-malware 2335caf4ec81ea00b6da46ec70fa9da3bd1887c3d87dd1fc24a9466f45f98e0e Any computer that has this package installed or running should be considered fully compromised. All secrets and keys stored on that computer should be...