The problem
Bots and automated scanners probe the infrastructure all day long. The bottleneck isn't generating more alerts. It's separating what actually deserves attention and responding while it still matters.
In this real case, the LLM only looks at traffic that got past the WAF. Anything the traditional control already blocked drops out before the pipeline.
The actual architecture
WAFEvery request becomes an event: IP, path, user-agent, status and action.
Collect + filterA Worker keeps a window of events and drops what was already blocked, known assets and predictable noise.
LLMReceives a structured summary and returns a decision as JSON.
GuardrailsAllow list, minimum confidence, exceptions, and actions enabled per category.
ActionA block with a TTL at the edge, plus an alert with evidence for review.
What goes into the AI
- A system prompt with a blue team analyst role, the real stack, and rules for what should not be reported.
- Top IPs, paths and user-agents for the period, always with HTTP status and ASN.
- Rare paths, where directory scanners and enumeration tend to show up.
- Suspects pre-flagged by simple rules, such as traversal, SQLi, .env and web shells, for the AI to validate.
- Operational context: cloud origin, repeat offenders, and block history.
What comes out of the AI
Free text does not drive action. The output is structured so code can validate it before anything executes.
{
"clean": false,
"threats": [{
"tipo": "Sensitive file probing",
"categoria": "git_env_exposto",
"host_path": "api.example.com/.env",
"ip": "203.0.113.7",
"severidade": "ALTO",
"confianca": 88,
"regra_sugerida": "Block /.env* at the WAF"
}]
}The category comes from a closed list. Severity and confidence are validated in code. If the model breaks the contract, it falls through to a fallback, never to an action.
Guardrails that make autonomy possible
Nothing blocks by defaultAutomated action has to be enabled per severity and category, with a defined list and TTL.
The allow list is absoluteCustomers, partners, employees and protected origins are never blocked, even on a detection.
Minimum confidenceBelow the cutoff, the system alerts and leaves the decision to a person.
Ambiguity brings in a humanWhen the context isn't enough, automated action is not the fallback.
Reversible actionBlocks expire by TTL, and every action leaves evidence behind.
Recipe for reproducing the idea
Ingredients
- WAF/CDN logs reachable via API, Logpush, a bucket or equivalent.
- A scheduled job, Worker, Lambda or serverless process.
- An LLM that can return structured JSON.
- A block list the WAF consumes, with a TTL.
- An alert channel with a link to the evidence.
Lessons that save weeks
- Pre-filter before the LLM. Predictable noise is better handled in code.
- Rules for the binary calls, AI for contextual judgment.
- HTTP status and context change what an event means.
- A false positive is a bug to fix, not a reason to turn the engine off.
- Start in recommendation mode. Raise autonomy once you've measured confidence and error.
Takeaway: AI isn't a report generator. It acts, within limits we set.
Take this case with you
In Markdown to paste into your AI, or as the starting point for your own playbook.