About this walkthrough
Every screen below is a real screenshot from an automated, end-to-end run of Regisseur against a live instance of the product, not a mock-up. Woodgrove Life is a fictional demo company; the people and cases shown are illustrative. Where something is set up for the demo, we say so on the screen it affects: the safety-limit test uses a disposable case with a deliberately low limit and pre-filled applicant details; the "human" calibration labels are entered by the test script from a fixed rule; the emergency brake is pulled and released within seconds. Some results are honest "not yet" states: the drift check refuses to call a trend, the AI judge is not calibrated so rubric scores read "—", the audit panel's event preview is below the fold in its screenshot, and the safety limit is not exact under parallel work. We show them as they came out.
The configuration screen for one agent, the Medical Review Prep Agent, which reads lab and physician records and is instructed never to make an underwriting decision itself. The screen shows its instructions and the three autonomy options (Always Review, Graduated, Always Autonomous); here "Always Autonomous" is selected, and the rationale field notes that this is forced for demo fidelity on an all-autonomous demo template. In a production setup you would normally choose a more cautious tier for a regulated task.
The agent's Testing tab before any run: three built-in cases and the first two custom cases (Complete-Clean Evidence, Missing APS) are in frame, with more below the fold. Each row shows how many checks, rubric questions and samples it runs. The suite exists before the agent touches a live case. (The banner at the top reads "Earlier version — not in use": this page was opened on an earlier saved version of the agent, and we have not hidden that.)
The case list after a new test case, "Elevated Blood Pressure — APS-Confirmed Hypertension," was entered through the Add test case form during this run. It now appears as the last custom row (1 check, 3 rubric questions, 2 samples). The form itself is not in this frame; only the saved result. The demo script deletes and re-creates this case each run so the suite does not pile up duplicates.
The model picker lists six model connections; two are ticked for this comparison, Claude Haiku 4.5 (the workspace default) and DeepSeek V3.2. The picker reads "9 cases". Recent runs from earlier sessions are listed below it. These are connections the customer configures with their own cloud account; Regisseur does not bill for the model usage.
The run has been started: the Run button reads "Running…", the two model checkboxes are locked, and the Current run panel lists both model connections as Pending. This is a live job on the queue, not a canned response.
A finished run, opened from Recent runs, for the two models. Claude Haiku 4.5: deterministic 100%, rubric 100%, p95 latency 7.01s, $0.00276 per run, PASS and marked Recommended as the cheapest PASS. DeepSeek V3.2: deterministic 100%, rubric 95%, p95 12.66s, $0.00086 per run, verdict PARTIAL. Both show "drift: not_comparable". The rubric scores count because, when this run was scored, the AI judge was calibrated (κ 1.00 on 8 human labels, shown in the calibration scenes). All figures are what this run measured.
The two result cards, side by side: Haiku (Recommended, PASS, rubric 100%) and DeepSeek V3.2 (PARTIAL, rubric 95%, cheaper per run but slower), and the "Drift vs previous run" section beginning just below them, where both models read "not_comparable".
An AI "judge" scores the subjective parts of each response, and it has to earn trust against human labels. This is the Judge Calibration panel when this run reached it: "Calibrated", κ 1.00, agreement 100%, 8 labels, judge qwen.qwen3-32b-v1:0. Carve-out: those 8 labels were entered earlier by the demo script, using a fixed rule, not by a person. The status is what the product computed from them.
The calibration banner after two more labels were entered in this run: it now reads "Not yet calibrated", κ 0.62 (need 0.8), because one new label disagreed with the judge. The newly labeled items are lower on the page (next screen). Carve-out: the demo script, not a person, enters the "human" labels by a fixed rule over each output. The script removes the two labels it created when the run ends and recomputes the judge's record, so this reading is a mid-run snapshot and the next run starts from the same 8 labels.
Items near the last label submitted in this run: submitted cards show "Judge verdict revealed" with the judge's verdict and an "agree" or "disagree" marker (one shows disagree: the script's fixed-rule label differed from the judge), and the card below, not yet labeled, still shows the Pass / Fail / Submit buttons. The judge's answer appears only after the label is submitted. Carve-out: the demo script applied the label by a fixed rule, not a person.
A disposable test case on the all-autonomous term-life process, after its safety limit tripped. Two carve-outs, both set by a setup script before the run: (1) the case's automated-step limit was lowered to 2 by a direct database change (the system would normally allow 23), so it trips within the first burst of work; (2) the applicant's name and date of birth were pre-filled on the case, because the intake agent no longer pulls them from the records system. The graph is cropped to the nodes: Application Received, Application Intake and Paramedical Exam Ordered are complete; the four parallel steps after it (Lab Order Sent, APS Request Sent, Background Check Ordered, MVR Check Ordered) show "Ready", meaning they did NOT run; three steps show "Awaiting upload"; the rest are pending.
The case paused itself and the banner says why: "execution_cap_exceeded: count=3 max=2", with "Agent runs: 3 attempted / 2 allowed (the last was refused)" and a Reset & resume button. The limit of 2 is the demo's setup override. Once tripped, the case refuses every further agent run until an operator resets it, so nothing past the refusal is counted or executed. (An earlier version of this product let the parallel steps keep running after the trip and showed "Executions: 7 / 2"; that was a real defect, since fixed.) A database check taken after the run, listing the case's event timeline, shows no agent task started between the trip and the reset, about six minutes.
After the operator clicked "Reset & resume", the banner is gone and the case shows Active. The case summary reads that application intake and paramedical exam ordering are complete and that it is waiting on lab results, APS records, background check and MVR check.
The same case's graph after the resume, cropped to the nodes: Lab Order Sent and APS Request Sent now show Complete, while Background Check Ordered and MVR Check Ordered still show Ready. The reset released the held work, but the demo's limit of 2 is still in force, so after two more runs the third attempt tripped the breaker again (confirmed by a database check of the event timeline); the remaining two steps stay held. Which two of the four steps run first can differ between runs.
Settings, Safety: the Emergency Brake, which halts automated agent tasks across the whole workspace. Shown in its normal state, with the workspace Status reading Active.
Pulling the brake opens a confirmation: a reason field and a box where HALT must be typed before "Confirm — Halt Now" is used. The demo reason says it will resume immediately after the capture.
With the brake engaged, a red banner reads "Emergency Brake is active — all automated actions are paused" across the product; here on the Operations Dashboard, a different screen from where the brake was pulled. (The dashboard also lists other demo cases from this workspace; this walkthrough does not use them.)
The same dashboard after the brake was released through the product's API: the banner is gone. A separate check of the halt status confirmed halted = false; that check is not visible on this screen.
The Safety & Governance Log on the Ops page, newest first: Emergency Brake released, Emergency Brake activated, Circuit breaker tripped (the test case, second trip after the reset), Circuit breaker reset, Circuit breaker tripped (first trip), then older brake events. The brake used in this run was released at the end; a database check after the run confirmed the workspace reads is_halted = false.
The case page scrolled to the bottom of the Audit Trail Export panel: a "Graph events" count (37 events), the covered time range, and the Download JSON and Download CSV buttons. Above it the Case Activity Feed shows only the newest entries, and the red banner shows the safety limit tripped again after the reset (the demo cap is still 2). The trip and reset entries are not in this frame; they are in the downloaded record shown next. The workspace-level brake is not a per-case event, so it is not in this export.
The same case's complete exported history, downloaded and opened in the browser's JSON viewer. The safety-limit trip is visible as an entry of type CIRCUIT_BREAKER_TRIPPED with cap 2 and observed 3, and the matching CIRCUIT_BREAKER_RESET entry from the operator's click is visible just below it. The file's own header counts 37 events.
A run exported as a standalone certificate, top of the page: verdict PASS, a SHA-256 tamper-evidence fingerprint, the agent model (Haiku 4.5), the judge model (qwen.qwen3-32b-v1:0), judge agreement 100% against human labels, Cohen's κ 1.000 over 8 calibration labels, and a "calibrated at" time. The gold-set identifier reads "trust-governance-cleanup:…" because the demo's label-cleanup script recomputed the judge record just before this capture; the labels are the 8 scripted ones. The agent definition version still reads "not-pinned-phase-1" and the code revision "runtime", so the certificate does not pin the exact agent version. The warning banner further down the page is not in this frame.
Settings, LLM Model Configurations: six model connections, all on AWS Bedrock (including DeepSeek V3.2), one tagged Default and one tagged Judge. No key values are shown, only name, provider and model id, and a connection-test badge. All six read "Stale", last tested 6/23/2026, so the demo does not claim these connections were re-tested.
Settings, PHI/PII Field Masking: when enabled, values of fields whose names match the listed patterns (SSN, date of birth, diagnosis, medication and others) are SHA-256 hashed before being written. Shown as currently configured: the toggle is off. The demo does not switch it on.
Settings, Plan & Usage: this development workspace reads "No license record found… Usage limits are not enforced", with 22 process templates and 8 members counted, and an Add license button. The page says so rather than inventing a plan name.