1 matches found
ClaimMirage: When Self-Claims in Domain Names Change LLM Threat Judgments
Short claims such as not-phishing or official can change how a large language model LLM judges a domain name, without explicit prompt-injection commands. We study this manipulation as ClaimMirage: a name under inspection claims its own safety or approval. We analyze 622,080 judgments across 64...
5.8AI score
SaveExploits0
20