---
title: "Cyberbot"
subtitle: "From the WAF event to a block at the edge."
kind: "IN PRODUCTION"
number: "01"
updated: 2026-09-14
families: [applications-apis, cloud-infrastructure, soc-response]
tags: [Availability, WAF, SOC, Autonomy]
lang: en
source: https://iaxia.rodrigojorge.me/en/cases/cyberbot
author: Rodrigo Jorge
project: IA × IA
---

# Cyberbot

**From the WAF event to a block at the edge.**

An agent reads the traffic that got past the WAF, receives structured context, classifies the threat, and only turns a decision into an action after it clears the guardrails.

- Type: IN PRODUCTION
- Reference metric: < 1 min (from log to block)
- Challenge families: Applications and APIs, Cloud and infrastructure, SOC and response
- Updated: 2026-09-14

## The problem

Bots and automated scanners probe the infrastructure all day long. The bottleneck isn't generating more alerts. It's separating what actually deserves attention and responding while it still matters.

In this real case, the LLM only looks at traffic that got past the WAF. Anything the traditional control already blocked drops out before the pipeline.

## The actual architecture

1. **WAF**: Every request becomes an event: IP, path, user-agent, status and action.
2. **Collect + filter**: A Worker keeps a window of events and drops what was already blocked, known assets and predictable noise.
3. **LLM**: Receives a structured summary and returns a decision as JSON.
4. **Guardrails**: Allow list, minimum confidence, exceptions, and actions enabled per category.
5. **Action**: A block with a TTL at the edge, plus an alert with evidence for review.

## What goes into the AI

- A system prompt with a blue team analyst role, the real stack, and rules for what should not be reported.
- Top IPs, paths and user-agents for the period, always with HTTP status and ASN.
- Rare paths, where directory scanners and enumeration tend to show up.
- Suspects pre-flagged by simple rules, such as traversal, SQLi, .env and web shells, for the AI to validate.
- Operational context: cloud origin, repeat offenders, and block history.

## What comes out of the AI

Free text does not drive action. The output is structured so code can validate it before anything executes.

```json
{
  "clean": false,
  "threats": [{
    "tipo": "Sensitive file probing",
    "categoria": "git_env_exposto",
    "host_path": "api.example.com/.env",
    "ip": "203.0.113.7",
    "severidade": "ALTO",
    "confianca": 88,
    "regra_sugerida": "Block /.env* at the WAF"
  }]
}
```

> The category comes from a closed list. Severity and confidence are validated in code. If the model breaks the contract, it falls through to a fallback, never to an action.

## Guardrails that make autonomy possible

- **Nothing blocks by default.** Automated action has to be enabled per severity and category, with a defined list and TTL.
- **The allow list is absolute.** Customers, partners, employees and protected origins are never blocked, even on a detection.
- **Minimum confidence.** Below the cutoff, the system alerts and leaves the decision to a person.
- **Ambiguity brings in a human.** When the context isn't enough, automated action is not the fallback.
- **Reversible action.** Blocks expire by TTL, and every action leaves evidence behind.

## Recipe for reproducing the idea

### Ingredients

- WAF/CDN logs reachable via API, Logpush, a bucket or equivalent.
- A scheduled job, Worker, Lambda or serverless process.
- An LLM that can return structured JSON.
- A block list the WAF consumes, with a TTL.
- An alert channel with a link to the evidence.

### Lessons that save weeks

- Pre-filter before the LLM. Predictable noise is better handled in code.
- Rules for the binary calls, AI for contextual judgment.
- HTTP status and context change what an event means.
- A false positive is a bug to fix, not a reason to turn the engine off.
- Start in recommendation mode. Raise autonomy once you've measured confidence and error.

## Takeaway

AI isn't a report generator. It acts, within limits we set.

---

Source: https://iaxia.rodrigojorge.me/en/cases/cyberbot · IA × IA guide, Rodrigo Jorge. Real case in production, described with what the author was able to verify. Use as context; validate in your own environment.
