Regisseur · Technical Walkthrough

Why you can trust agents with regulated work

A real, live-captured walkthrough of the guardrails behind every AI agent in Regisseur: a standing test suite, real multi-model cost comparisons, a judge that has to earn calibration, drift checking, a safety limit tripped on purpose to prove it works, a workspace-wide emergency stop, a tamper-evident audit certificate, and full control over your own AI keys and data.

Woodgrove Life — a Regisseur demo workspace (fictional) 26 screens, all captured live from a real run
9
Test cases run against one agent
2
AI models compared, real cost measured
2
Independent safety limits, both tested for real
JSON/CSV
Exportable case history
26
Screens captured live

About this walkthrough

Every screen below is a real screenshot from an automated, end-to-end run of Regisseur against a live instance of the product, not a mock-up. Woodgrove Life is a fictional demo company; the people and cases shown are illustrative. Where something is set up for the demo, we say so on the screen it affects: the safety-limit test uses a disposable case with a deliberately low limit and pre-filled applicant details; the "human" calibration labels are entered by the test script from a fixed rule; the emergency brake is pulled and released within seconds. Some results are honest "not yet" states: the drift check refuses to call a trend, the AI judge is not calibrated so rubric scores read "—", the audit panel's event preview is below the fold in its screenshot, and the safety limit is not exact under parallel work. We show them as they came out.

Part 1
Every Agent Ships With a Test Suite
A real eval case authored live, a real multi-model spend comparison, and a judge that has to earn the right to be trusted
1
Meet the agent, before it does any work
The configuration screen for one agent, the Medical Review Prep Agent, which reads lab and physician records and is instructed never to make an underwriting decision itself. The screen shows its instructions and the three autonomy options (Always Review, Graduated, Always Autonomous); here "Always Autonomous" is selected, and the rationale field notes that this is forced for demo fidelity on an all-autonomous demo template. In a production setup you would normally choose a more cautious tier for a regulated task.
Meet the agent, before it does any work
2
A standing test suite, not a one-time check
The agent's Testing tab before any run: three built-in cases and the first two custom cases (Complete-Clean Evidence, Missing APS) are in frame, with more below the fold. Each row shows how many checks, rubric questions and samples it runs. The suite exists before the agent touches a live case. (The banner at the top reads "Earlier version — not in use": this page was opened on an earlier saved version of the agent, and we have not hidden that.)
A standing test suite, not a one-time check
3
Adding a new test case, live
The case list after a new test case, "Elevated Blood Pressure — APS-Confirmed Hypertension," was entered through the Add test case form during this run. It now appears as the last custom row (1 check, 3 rubric questions, 2 samples). The form itself is not in this frame; only the saved result. The demo script deletes and re-creates this case each run so the suite does not pile up duplicates.
Adding a new test case, live
4
Two of your own AI models, side by side
The model picker lists six model connections; two are ticked for this comparison, Claude Haiku 4.5 (the workspace default) and DeepSeek V3.2. The picker reads "9 cases". Recent runs from earlier sessions are listed below it. These are connections the customer configures with their own cloud account; Regisseur does not bill for the model usage.
Two of your own AI models, side by side
5
The comparison, running for real
The run has been started: the Run button reads "Running…", the two model checkboxes are locked, and the Current run panel lists both model connections as Pending. This is a live job on the queue, not a canned response.
The comparison, running for real
6
Real cost, real speed, real pass/fail, for each model
A finished run, opened from Recent runs, for the two models. Claude Haiku 4.5: deterministic 100%, rubric 100%, p95 latency 7.01s, $0.00276 per run, PASS and marked Recommended as the cheapest PASS. DeepSeek V3.2: deterministic 100%, rubric 95%, p95 12.66s, $0.00086 per run, verdict PARTIAL. Both show "drift: not_comparable". The rubric scores count because, when this run was scored, the AI judge was calibrated (κ 1.00 on 8 human labels, shown in the calibration scenes). All figures are what this run measured.
Real cost, real speed, real pass/fail, for each model
7
Which model should you actually use?
The two result cards, side by side: Haiku (Recommended, PASS, rubric 100%) and DeepSeek V3.2 (PARTIAL, rubric 95%, cheaper per run but slower), and the "Drift vs previous run" section beginning just below them, where both models read "not_comparable".
Which model should you actually use?
8
Before we trust the AI judge, we test the judge
An AI "judge" scores the subjective parts of each response, and it has to earn trust against human labels. This is the Judge Calibration panel when this run reached it: "Calibrated", κ 1.00, agreement 100%, 8 labels, judge qwen.qwen3-32b-v1:0. Carve-out: those 8 labels were entered earlier by the demo script, using a fixed rule, not by a person. The status is what the product computed from them.
Before we trust the AI judge, we test the judge
9
The judge is checked against new human labels
The calibration banner after two more labels were entered in this run: it now reads "Not yet calibrated", κ 0.62 (need 0.8), because one new label disagreed with the judge. The newly labeled items are lower on the page (next screen). Carve-out: the demo script, not a person, enters the "human" labels by a fixed rule over each output. The script removes the two labels it created when the run ends and recomputes the judge's record, so this reading is a mid-run snapshot and the next run starts from the same 8 labels.
The judge is checked against new human labels
10
Human verdict first, then the judge’s
Items near the last label submitted in this run: submitted cards show "Judge verdict revealed" with the judge's verdict and an "agree" or "disagree" marker (one shows disagree: the script's fixed-rule label differed from the judge), and the card below, not yet labeled, still shows the Pass / Fail / Submit buttons. The judge's answer appears only after the label is submitted. Carve-out: the demo script applied the label by a fixed rule, not a person.
Human verdict first, then the judge’s
Part 2
Catching Drift Before It Becomes a Problem
A run-over-run check that refuses to claim a trend it cannot support
11
Drift checking, and an honest “not comparable”
Each run is compared with the previous run of the same model. Both models read "drift: not_comparable" with the message "Prompt/pipeline changed since baseline — comparison refused, not a measured trend", and every metric is tagged "unmeasured"; old and new values are listed but no verdict is claimed. A PASS/WARN/FAIL drift verdict needs two runs on the same agent definition. We show the honest "not yet" rather than a staged regression.
Drift checking, and an honest “not comparable”
Part 3
The Runaway That Wasn’t
A per-case safety limit, tripped for real, with a bigger switch above it
12
A safety limit, tripped on purpose to prove it works
A disposable test case on the all-autonomous term-life process, after its safety limit tripped. Two carve-outs, both set by a setup script before the run: (1) the case's automated-step limit was lowered to 2 by a direct database change (the system would normally allow 23), so it trips within the first burst of work; (2) the applicant's name and date of birth were pre-filled on the case, because the intake agent no longer pulls them from the records system. The graph is cropped to the nodes: Application Received, Application Intake and Paramedical Exam Ordered are complete; the four parallel steps after it (Lab Order Sent, APS Request Sent, Background Check Ordered, MVR Check Ordered) show "Ready", meaning they did NOT run; three steps show "Awaiting upload"; the rest are pending.
A safety limit, tripped on purpose to prove it works
13
Paused, with the reason stated plainly
The case paused itself and the banner says why: "execution_cap_exceeded: count=3 max=2", with "Agent runs: 3 attempted / 2 allowed (the last was refused)" and a Reset & resume button. The limit of 2 is the demo's setup override. Once tripped, the case refuses every further agent run until an operator resets it, so nothing past the refusal is counted or executed. (An earlier version of this product let the parallel steps keep running after the trip and showed "Executions: 7 / 2"; that was a real defect, since fixed.) A database check taken after the run, listing the case's event timeline, shows no agent task started between the trip and the reset, about six minutes.
Paused, with the reason stated plainly
14
One click to resume, once reviewed
After the operator clicked "Reset & resume", the banner is gone and the case shows Active. The case summary reads that application intake and paramedical exam ordering are complete and that it is waiting on lab results, APS records, background check and MVR check.
One click to resume, once reviewed
15
Back in motion
The same case's graph after the resume, cropped to the nodes: Lab Order Sent and APS Request Sent now show Complete, while Background Check Ordered and MVR Check Ordered still show Ready. The reset released the held work, but the demo's limit of 2 is still in force, so after two more runs the third attempt tripped the breaker again (confirmed by a database check of the event timeline); the remaining two steps stay held. Which two of the four steps run first can differ between runs.
Back in motion
16
The bigger switch: an emergency stop for everything
Settings, Safety: the Emergency Brake, which halts automated agent tasks across the whole workspace. Shown in its normal state, with the workspace Status reading Active.
The bigger switch: an emergency stop for everything
17
No accidental "stop everything"
Pulling the brake opens a confirmation: a reason field and a box where HALT must be typed before "Confirm — Halt Now" is used. The demo reason says it will resume immediately after the capture.
No accidental "stop everything"
18
Impossible to miss, once it’s pulled
With the brake engaged, a red banner reads "Emergency Brake is active — all automated actions are paused" across the product; here on the Operations Dashboard, a different screen from where the brake was pulled. (The dashboard also lists other demo cases from this workspace; this walkthrough does not use them.)
Impossible to miss, once it’s pulled
19
Released
The same dashboard after the brake was released through the product's API: the banner is gone. A separate check of the halt status confirmed halted = false; that check is not visible on this screen.
Released
20
One log for every safety action
The Safety & Governance Log on the Ops page, newest first: Emergency Brake released, Emergency Brake activated, Circuit breaker tripped (the test case, second trip after the reset), Circuit breaker reset, Circuit breaker tripped (first trip), then older brake events. The brake used in this run was released at the end; a database check after the run confirmed the workspace reads is_halted = false.
One log for every safety action
Part 4
Every Action, On the Record
A timestamped case history, and a tamper-evident certificate that is honest about its own limits
21
A timestamped history for every case
The case page scrolled to the bottom of the Audit Trail Export panel: a "Graph events" count (37 events), the covered time range, and the Download JSON and Download CSV buttons. Above it the Case Activity Feed shows only the newest entries, and the red banner shows the safety limit tripped again after the reset (the demo cap is still 2). The trip and reset entries are not in this frame; they are in the downloaded record shown next. The workspace-level brake is not a per-case event, so it is not in this export.
A timestamped history for every case
22
The full record, not just the preview
The same case's complete exported history, downloaded and opened in the browser's JSON viewer. The safety-limit trip is visible as an entry of type CIRCUIT_BREAKER_TRIPPED with cap 2 and observed 3, and the matching CIRCUIT_BREAKER_RESET entry from the operator's click is visible just below it. The file's own header counts 37 events.
The full record, not just the preview
23
A certificate that can’t be quietly edited, and says what it doesn’t pin
A run exported as a standalone certificate, top of the page: verdict PASS, a SHA-256 tamper-evidence fingerprint, the agent model (Haiku 4.5), the judge model (qwen.qwen3-32b-v1:0), judge agreement 100% against human labels, Cohen's κ 1.000 over 8 calibration labels, and a "calibrated at" time. The gold-set identifier reads "trust-governance-cleanup:…" because the demo's label-cleanup script recomputed the judge record just before this capture; the labels are the 8 scripted ones. The agent definition version still reads "not-pinned-phase-1" and the code revision "runtime", so the certificate does not pin the exact agent version. The warning banner further down the page is not in this frame.
A certificate that can’t be quietly edited, and says what it doesn’t pin
Part 5
Your Keys. Your Data.
Every model credential is yours; PHI masking and license status are one click away
24
Every AI connection is yours
Settings, LLM Model Configurations: six model connections, all on AWS Bedrock (including DeepSeek V3.2), one tagged Default and one tagged Judge. No key values are shown, only name, provider and model id, and a connection-test badge. All six read "Stale", last tested 6/23/2026, so the demo does not claim these connections were re-tested.
Every AI connection is yours
25
Sensitive-field protection, off by default
Settings, PHI/PII Field Masking: when enabled, values of fields whose names match the listed patterns (SSN, date of birth, diagnosis, medication and others) are SHA-256 hashed before being written. Shown as currently configured: the toggle is off. The demo does not switch it on.
Sensitive-field protection, off by default
26
Licensing, stated plainly
Settings, Plan & Usage: this development workspace reads "No license record found… Usage limits are not enforced", with 22 process templates and 8 members counted, and an Add license button. The page says so rather than inventing a plan name.
Licensing, stated plainly