Post
1242
Will your AI agent tell you it was attacked?
We took the same agent from our earlier experiment and added one thing: a twentieth tool,
The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models.
Alarm rates ranged from 49% to zero.
The unexpected result came from the newest model in the test,
Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished.
That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability.
A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it.
Full experiment and results:
https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked
Quadrat-IPI dataset:
mihailgribov/quadrat-ipi
Run your own model:
https://github.com/mihail-gribov/quadrat-ipi-model-eval
#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents
We took the same agent from our earlier experiment and added one thing: a twentieth tool,
escalate_security_incident.The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models.
Alarm rates ranged from 49% to zero.
The unexpected result came from the newest model in the test,
gpt-6-astra.Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished.
That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability.
A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it.
Full experiment and results:
https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked
Quadrat-IPI dataset:
mihailgribov/quadrat-ipi
Run your own model:
https://github.com/mihail-gribov/quadrat-ipi-model-eval
#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents