Lucene search
+L

2 matches found

Kitploit
Kitploit
•added 2026/10/06 3:59 a.m.•3 views

benign-instruction-bench

훈련한 테스트를 통과하기 프롬프트 인젝션 탐지기를 LLM 에이전트가 실제로 사용하는 지점, 즉 에이전트가 읽는 도구 출력에서 재평가한다. 팀들은 벤치마크 점수로 인젝션 탐지기를 고른다. 우리는 그 점수가 에이전트 내부에서의 동작을 예측하는지 확인하고, 대체로 그 점수는 벤치마크가 탐지기의 훈련 데이터에 얼마나 가까운지를 측정한다는 것을 발견한다. 발견 사항 15개의 탐지기오픈 9개, Meta의 Prompt Guard 2를 포함한 라이선스 제한 6개와 2개의 작업 인식 LLM 판정자를, 두 개의 에이전트 벤치마크AgentDojo...

6.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/11 12:00 a.m.•11 views

Effective Red-Teaming of Policy-Adherent Agents

Task-oriented LLM-based agents are increasingly used in domains with strict policies, such as refund eligibility or cancellation rules. The challenge lies in ensuring that the agent consistently adheres to these rules and policies, appropriately refusing any request that would violate them, while...

6.9AI score
SaveExploits0
Rows per page
Query Builder