The problem
An autonomous agent can query SIEM, EDR, IAM, WAF and cloud, correlate signals, and arrive at a solid hypothesis. The risk shows up when the same entity that investigates also decides, on its own, how far it can act.
The proposal is to separate roles and authority. The operator stays powerful for investigation. Authorization goes through a layer with a much narrower objective, less context, fewer tools and less authority.
Architecture
The contract between operator and supervisor
The operator doesn't hand the supervisor an entire conversation. It hands over a structured request with action, target, evidence, confidence and reversibility.
{
"action": "revoke_session",
"target": "user_8271",
"evidence": {
"credential_leak": "confirmed",
"impossible_travel": true,
"active_sessions": 3
},
"operator_confidence": 94,
"reversible": true
}The response is a contract too
{
"decision": "APPROVE",
"policy": "IAM-07",
"scope": "user_8271",
"ttl": "15m",
"requires_human": false
}Same incident. Different authority.
The authority matrix
Can be automated
- Enrich and correlate alerts.
- Build a timeline and gather evidence.
- Temporarily block an IP within policy.
- Revoke a session when identity, scope and reversibility are within the rule.
Escalates to a supervisor
- Disabling a privileged account.
- Changing a critical configuration.
- Any action with a large blast radius.
- Any irreversible or critical action.
Authority that is earned and lost
The authority to act without approval isn't a single level for the agent, and it isn't permanent. It is granted per action class, after the agent's decisions in that class have been compared against human review on a defined sample. The agent doesn't measure its own accuracy.
When verdicts start diverging from independent review beyond a threshold, the class drops back to its previous mode. The threshold, the observation window and who reviews are part of the contract, not left implicit.
The agent badge
What the badge declares
- The agent's identity and its human owner.
- The context it can query and the access it holds.
- Authorized actions, per class, with each one's current level.
- Required supervision, required evidence, and validity with an expiration.
What stays off the badge
- The kill switch. The agent doesn't control its own shutdown.
- Class demotion. The platform executes it.
- The accuracy measurement. It comes from independent review.
- An agent without an identity and bounded authority is a service account with too much power.
What actually reduces the risk
A smaller model doesn't mean a safer model
A smaller model can work well as a supervisor because the task is narrow, but size is no guarantee of safety. The control comes from the architecture: limited role, minimal context, strict contract, less authority, independence and deterministic validation.
A second identical instance doesn't solve the problem on its own either. If the operator and the supervisor share the same assumptions, context and way of deciding, the second layer can repeat the same failure.
How to implement it
In Markdown to paste into your AI, or as the starting point for your own playbook.