---
title: "Agent supervisor"
subtitle: "More intelligence in the operator. Less power in whoever authorizes."
kind: "PLAYBOOK"
number: "07"
updated: 2026-10-06
families: [ai-agent-governance, soc-response, identity-access]
tags: [Agents, Governance, Autonomy, Policy Engine]
lang: en
source: https://iaxia.rodrigojorge.me/en/cases/supervisor-agentes
author: Rodrigo Jorge
project: IA × IA
---

# Agent supervisor

**More intelligence in the operator. Less power in whoever authorizes.**

An operator agent investigates with broad context and proposes an action. A deliberately limited supervisor evaluates the request. A Policy Engine applies the objective rules before anything executes.

- Type: PLAYBOOK
- Reference metric: 3 layers (investigate, authorize and execute)
- Challenge families: AI and agent governance, SOC and response, Identity and access
- Updated: 2026-10-06

## The problem

An autonomous agent can query SIEM, EDR, IAM, WAF and cloud, correlate signals, and arrive at a solid hypothesis. The risk shows up when the same entity that investigates also decides, on its own, how far it can act.

The proposal is to separate roles and authority. The operator stays powerful for investigation. Authorization goes through a layer with a much narrower objective, less context, fewer tools and less authority.

## Architecture

1. **Operator agent**: The more capable model. Investigates, correlates, builds a hypothesis, gathers evidence and requests an action.
2. **Supervisor agent**: Receives only the context needed for one specific decision. Approves, denies or escalates.
3. **Policy Engine**: Applies deterministic rules: allow and deny lists, authority, asset, identity, reversibility, TTL, schema and kill switch.
4. **Action**: Executes only within the delegated authority. Outside it, a human is required.

> The supervisor doesn't need to investigate better than the operator. It needs to decide something much narrower.

## The contract between operator and supervisor

The operator doesn't hand the supervisor an entire conversation. It hands over a structured request with action, target, evidence, confidence and reversibility.

```json
{
  "action": "revoke_session",
  "target": "user_8271",
  "evidence": {
    "credential_leak": "confirmed",
    "impossible_travel": true,
    "active_sessions": 3
  },
  "operator_confidence": 94,
  "reversible": true
}
```

## The response is a contract too

```json
{
  "decision": "APPROVE",
  "policy": "IAM-07",
  "scope": "user_8271",
  "ttl": "15m",
  "requires_human": false
}
```

> Free text can explain. The authorization that moves on to the next layer has to respect a verifiable schema.

## Same incident. Different authority.

**Operational example**

| | |
|---|---|
| Request 1 | Block a malicious IP for 30 minutes |
| Supervisor | APPROVE |
| Policy Engine | Temporary, reversible action within the delegated authority |
| Outcome | Executes automatically |
| Request 2 | Disable the CFO's account |
| Supervisor | May consider the evidence sufficient |
| Policy Engine | Privileged identity and a high-impact action |
| Outcome | HUMAN APPROVAL REQUIRED |

## The authority matrix

### Can be automated

- Enrich and correlate alerts.
- Build a timeline and gather evidence.
- Temporarily block an IP within policy.
- Revoke a session when identity, scope and reversibility are within the rule.

### Escalates to a supervisor

- Disabling a privileged account.
- Changing a critical configuration.
- Any action with a large blast radius.
- Any irreversible or critical action.

## Authority that is earned and lost

The authority to act without approval isn't a single level for the agent, and it isn't permanent. It is granted per action class, after the agent's decisions in that class have been compared against human review on a defined sample. The agent doesn't measure its own accuracy.

When verdicts start diverging from independent review beyond a threshold, the class drops back to its previous mode. The threshold, the observation window and who reviews are part of the contract, not left implicit.

- **Granted per class.** Enriching, blocking an IP for a limited time and revoking a session are separate classes. Each one levels up on its own, on its own evidence.
- **Independent evaluation.** The sample of decisions is reviewed by someone who is neither the operator nor the supervisor. The agent doesn't grade its own exam.
- **Demotion trigger.** Divergence beyond the threshold demotes the class automatically. The demotion is executed by the platform, outside the model.
- **Actions in flight.** The contract spells out what happens to what was already executed when a class drops a level: rollback, compensation, or just a record. Expiring the authorization doesn't undo what's already done.

## The agent badge

### What the badge declares

- The agent's identity and its human owner.
- The context it can query and the access it holds.
- Authorized actions, per class, with each one's current level.
- Required supervision, required evidence, and validity with an expiration.

### What stays off the badge

- The kill switch. The agent doesn't control its own shutdown.
- Class demotion. The platform executes it.
- The accuracy measurement. It comes from independent review.
- An agent without an identity and bounded authority is a service account with too much power.

## What actually reduces the risk

- **Narrow role.** The supervisor doesn't get an open-ended mission to investigate the whole incident. It evaluates one specific request.
- **Minimal context.** It receives only the data needed to decide, which shrinks the surface for interpretation.
- **Less authority.** It doesn't get free access to the environment and doesn't execute the requested action directly.
- **Independence.** Objective, context and contract should keep it from simply replaying the operator's chain of reasoning.
- **Deterministic control.** Objective policies stay outside the LLM whenever they can be expressed as a rule.
- **Human by impact.** The human comes in by exception, by policy or by impact. Anything irreversible or critical requires human approval.

## A smaller model doesn't mean a safer model

A smaller model can work well as a supervisor because the task is narrow, but size is no guarantee of safety. The control comes from the architecture: limited role, minimal context, strict contract, less authority, independence and deterministic validation.

A second identical instance doesn't solve the problem on its own either. If the operator and the supervisor share the same assumptions, context and way of deciding, the second layer can repeat the same failure.

## How to implement it

1. **1. Catalog the actions**: List what your agents can request: query, block, revoke, isolate, change or delete.
2. **2. Classify impact**: Define reversibility, criticality, blast radius, and protected identities or assets.
3. **3. Define authority**: Map each action to automatic, supervisor, or human approval.
4. **4. Create contracts**: Standardize request and response with a verifiable schema and required evidence.
5. **5. Pull rules out of the LLM**: Allow list, limits, TTL, scope, privileged identities and kill switch live in code or in a policy engine.
6. **6. Log everything**: Request, evidence, decision, policy applied, execution, outcome, and any human intervention.

## Takeaway

We're not setting up one AI to trust another AI. We're separating authority.

---

Source: https://iaxia.rodrigojorge.me/en/cases/supervisor-agentes · IA × IA guide, Rodrigo Jorge. Defense playbook: an architecture to adapt, not evidence from production. Use as context; validate in your own environment.
