There is a quick way to make a defensive agent get things wrong: tell it to "block attacks". It reads like an instruction. In practice it is a question that is far too broad, and the model answers it the way it answers any broad question: with a plausible, confident opinion that is hard to audit.
The problem is not a weak model. It is a wide question. "Is this an attack?" requires the model to know, all at once, what SQL injection is, what XSS is, what path traversal, credential stuffing, scraping, a legitimate monitoring bot and a customer running an old plugin look like. Each of those has different signals, different false positives and different consequences when the call is wrong. Thrown into one bucket they become an average, and an average does not block anything well.
Narrow question, better answer
What works is the opposite: split by category and ask one question at a time. Instead of "is this an attack?", the agent gets "is this SQL injection?", with a prompt written for SQL injection, examples of SQL injection, examples of what is not SQL injection, and the confidence threshold the operation accepts for SQL injection. Then the same for XSS, for traversal, for every category the organization has decided to handle.
The difference shows up in three places:
- Precision. The model is right more often when the answer space is small and the examples come from the same family.
- Calibration. Each category gets its own threshold. The confidence that is enough to block an obvious traversal is not the confidence you want for a credential-stuffing pattern, where blocking a legitimate customer is expensive.
- Correction. When the AI gets it wrong, you know which category it got wrong, and you adjust that prompt, those examples or that threshold without touching the rest. A false positive becomes a localized bug, not a reason to switch the engine off.
How this shows up in Cyberbot
The Cyberbot case, running in production, does exactly this, and that is how it got to under a minute from log to block without taking customers down. The category comes from a closed list; the model does not invent new ones. Severity and confidence are validated by code. Simple rules pre-detect the obvious suspects (traversal, SQLi, exposed .env, web shells) and hand the AI only what needs contextual judgment. And automatic action is enabled per category and severity, with a defined list and TTL, never wholesale.
If the model steps outside the contract and returns something off the list, the request falls through to a fallback. Nothing happens. That is the part that matters most: the narrow question is not a prompt trick, it is what makes it possible to put a verifiable contract between the model and the action.
What to change in your playbook
If you are designing an agent that is allowed to act, the rule fits in one line: no action is authorized for "attack". Actions are authorized per category, and each category has its own prompt, examples and threshold, reviewed separately.
The playbook generator in this guide now includes that in its list of guardrails. It is not an implementation detail. It is the difference between an agent the operation trusts and one the operation turns off in its first week.