Every agent runs on zero standing trust, bounded before it acts, watched while it acts, and provable after. You don’t trust the AI; you bound it with deterministic boundaries, then replay exactly what it did. Evals right-size each step; when something fails, recovery is proposed and human-approved, never silent.
§ 01 · The trust lifecycle
Govern it before. Watch it during. Verify it after.
01Before it runsBind it
No agent gets standing trust.
Before an agent touches real work it is bounded, scoped to which tools it can call, which data it can read, and how far it can act without a human. Permissions are least-privilege and explicit, the autonomy ceiling is set by operations (the agent can’t self-promote), and the model behind each step is the cheapest one that passed its evals, not a default.
A platform coordinator tracks every case, surfaces work that has stalled, and chases the outside parties a case is waiting on. Every step carries an SLA clock with escalation. When a step fails, recovery routes are scored and proposed, but a person authorizes the one that ships. Two brakes sit above it all.
Coordinator
Surfaces stalls and stuck nodes before they breach
Every cognitive step writes to an append-only ledger with a propagating trace_id. Reconstruct exactly what happened (who decided, which rule version ran, what evidence was used) and re-run a past case on the forms, agents, and rules that were live when it opened. This is the half most agent stacks can’t hand you.
Append-only ledger
Every decision recorded, immutable (S.01)
Attribution
Which agent or human did what, and why
Deterministic replay
Reconstruct the case from its events
Regulatory versioning
Past work re-runs on the rules that were live (S.02)
The audit trailHuman and agent actions, interleaved and attributed, on one case record
We’re not selling you a Zero Trust Architecture. That’s a network and identity stack. We took its one durable idea (never trust, always verify) and applied it to the newest actors in your operation: AI agents, treated as non-human identities.
Least privilege
Scoped tokens, bounded tools, and autonomy ceilings, every grant explicit and revocable
No implicit trust
Regulated actions require human approval; an agent is never trusted by default
Always verify
An append-only, replayable audit trail with per-action attribution
Identity-aware
Every agent is a non-human identity with logged, bounded permissions
§ 03 · The seven structural guarantees
Each phase rests on a platform-level mechanism.
Seven of them. Every guarantee below exists whether or not an operator remembers it, enforced by the engine, not a policy document.
S.01Immutability
Audit trail on every decision.
Append-only event ledger with a propagating trace_id on every cognitive step. Reconstruct what happened, who decided, which rule version ran, and what evidence was used.
Past work remains re-runnable on the exact forms, agents, and rules that were live when it opened. Mid-flight changes do not rewrite running workloads.
Version model
SCD Type 2
Resolution
At case instantiation
Replay fidelity
Byte-identical
S.03Earned autonomy
Configurable autonomy ceilings.
Every agent operates within an operations-set ceiling. The engine enforces it, the agent cannot self-promote, and every promotion is logged.
Level 1 · Default
Always review
Every action routes to a human reviewer before execution. No agent ships without this first.
Level 2 · Earned
Graduated
Routine actions execute; flagged or low-confidence outputs halt for human sign-off. Most agents operate here.
Level 3 · Restricted
Fully autonomous
Reserved for mechanical tasks only. Scoped workflows; never PHI, never pricing, never signed outputs.
S.04Structural determinism
Deterministic, typed agent pipelines.
Agent execution is a deterministic pipeline. AI is one bounded step with typed input and output; validation, MCP calls, schema checks, and routing remain explicit.
Pipeline model
Typed DAG
LLM surface
Single step per node
Contract
Zod + JSON schema
S.05Halt semantics
Workspace-wide emergency brake.
One action pauses every agent, queue, and outbound message across a workspace. A halted workspace resumes only on explicit operator release. Below it, a per-job circuit breaker trips a single runaway case, without halting the workspace.
Scope
Workspace-wide
Enforcement
Every task, every tick
Release
Explicit operator command
S.06Operator-owned
Configurable without engineers.
Process managers adjust graphs, prompts, autonomy, and MCP scopes in conversation. Every change is versioned; engineering owns the engine, not the ops logic.
Authoring
Conversational graph
Change unit
Versioned template
Deploy
Zero · published
S.07Rehearsed
Tested before trusted.
Run any agent against a real scenario with the outside world mocked (no email sent, no record written, no external call) while its reasoning runs for real, and inspect every step of the trajectory before it goes live. Optionally gate publishing on a passing eval: a draft cannot ship unless its latest eval run passed and that run is pinned to the exact draft being published.
Run mode
Dry-run · side effects mocked
Inspection
Full step trajectory
Publish gate
Opt-in · fail-closed · pinned
§ 04 · Why it’s different
RPA can’t reason. Agent frameworks can’t account for themselves.
Legacy RPA has no judgment; raw agent frameworks have no causal audit, no runtime compliance, no autonomy ceiling, and no brake. Regisseur is the production scaffold that makes an agent both capable and accountable, the difference between a demo and something legal will sign off.
We’ll bound an agent, run a case, trip the emergency brake, and then replay the whole trace (who decided what, on which rule version, with which evidence).