6
Test scenarios behind one agent
2
AI models compared, real cost measured
2
Independent safety limits, both tested for real
41
Timestamped events on one case, exportable
About this walkthrough
Every screen below is a real screenshot from an automated, end-to-end run of Regisseur against a live instance of the product — not a mock-up, not a slide deck. Woodgrove Life is a fictional demo company; the people and cases shown are illustrative. Running this walkthrough as a real end-to-end test — not writing it as a script — has, over successive rehearsals, surfaced a few genuine gaps in how the product displayed information: a governance-tier setting that wasn't shown on screen, a case-history record that was briefly missing, and a workspace-wide safety log that was only showing half of what it should have. We caption each of those honestly, in place, rather than editing around them — and each was found and fixed the same day, with the screenshots you're looking at now showing the fixed state. We would rather earn your trust with an honest demo than lose it with a polished one.
This is the configuration screen for one agent — the Medical Review Prep Agent, which reads lab and physician records and is explicitly instructed never to make an underwriting decision itself. Below its instructions sit three governance-tier options for how much this agent can do without a human in the loop; here it's set to the most autonomous tier. An earlier run of this exact demo caught this same screen failing to show which tier was active — a real display gap we found by testing our own product. We fixed it, and this screenshot is the proof: run it again, and it now shows correctly.
Before any new work happens, this agent already has five test scenarios sitting behind it, added by an administrator in earlier working sessions — clean evidence, missing records, a flagged biomarker, sparse labs, and a conflicting-history check. No agent goes live without a test suite. This is what that looks like.
A sixth test case, "Elevated Blood Pressure — APS-Confirmed Hypertension," entered directly into this screen — no code, no engineering ticket. Any administrator with access can extend an agent's test coverage as their real-world cases surface new scenarios.
Two model connections selected for comparison — Claude Haiku (this workspace's everyday default) and DeepSeek (a second option worth comparing). Both are connections our customer configured with their own account and their own credentials; we never see or bill for the underlying AI usage.
Both models are now actually being tested against every scenario in the suite — real requests, in progress. Nothing here is a canned demo response; this is the same comparison an administrator would run before trusting a new model with real casework.
The finished comparison: actual dollar cost per run, actual response time, and a pass/fail verdict against every check in the suite, for each model. This is measured spend on the customer's own account, not an estimate.
A side-by-side recommendation — the cheapest model that still passed every check, clearly marked, next to a full breakdown of how each model performed against the last time this same comparison ran. This is a data-backed recommendation, not a sales pitch for either AI vendor.
A second layer of quality control: an AI "judge" model scores the more subjective parts of each response, but that judge doesn't get trusted automatically. Here, four real human-reviewed samples have been collected so far — short of the five we require before we call a judge calibrated. We show that honestly rather than rounding up.
Two more real reviews, entered blind — meaning the reviewer commits to a verdict before ever seeing what the AI judge decided, so the human judgment can't be swayed by the AI's opinion. That brought the running total to six, past our five-review minimum, and the judge model agreed with the human reviewer on every single one: "Calibrated ✓, 100% agreement, 6 labels." We only start trusting an AI judge's scoring after it has proven itself against real human review — and here, it just did, live.
To prove our per-case safety limit works, we deliberately set an unusually low ceiling on how many automated steps one case could take before pausing — far lower than the system would normally allow — on a disposable test case created just for this. The workflow ran several steps in parallel and the limit caught it, exactly as designed.
The case paused itself and says exactly why: it hit its execution limit. Nothing was lost — the case is paused, not broken, and an operator can review and resume it with one click. We also want to be candid about something this test revealed: when several steps start at almost the same moment, a couple can slip through just after the limit is hit before the pause takes full effect, even though every later step correctly stays blocked. We're tightening that timing window; it does not affect the intended safety limit, which held.
An operator reviewed the pause and clicked "Reset & resume." The case picks back up immediately from exactly where it left off — no lost work, no do-over.
The same case, now continuing normally after the operator's review. A safety pause is a checkpoint, not a failure.
Above the per-case safety limit sits a single, bigger control: an emergency brake that can pause every automated task across the entire workspace at once, for situations the per-case limit isn't meant to catch. Shown here in its normal, off state.
Pulling the emergency brake requires typing the word "HALT" to confirm, with an optional reason recorded for the record. This is a deliberate, hard-to-trigger-by-accident action.
With the brake engaged, a plain warning appears across the entire product for every user — here, on the main operations dashboard, a different screen from where the brake was pulled. Every automated task checks for this before doing anything and stands down immediately if it's on.
The same dashboard, moments later, with the banner gone. We verified directly that the workspace was genuinely back to normal before continuing — nothing was silently left in a paused state. Every case that was held during the pause resumes exactly where it left off.
A single, workspace-wide log listing every time an automatic safety limit or a manual stop fired, across every case, in one place. Here it correctly shows both halves at once: the broader manual stop demonstrated on the previous screens — activated, then released, moments ago — sitting right alongside this test's own automatic safety-limit trip and reset from earlier in the walkthrough. An earlier run of this exact test caught the manual-stop half missing from this same log entirely, a real permissions gap in how that one event type was stored. We found it, fixed it the same day, and this screenshot is the proof: run the test again, and both halves now show correctly.
Every case carries its own exportable history — 41 timestamped entries here, downloadable as JSON or CSV, built for handing to an auditor. The top entry in this on-screen preview is the exact moment this case's own safety limit was reset, timestamped and attributed to a named admin — an earlier run of this demo caught this specific record missing entirely, and it's now genuinely part of the case's own trail. One honest caveat: this preview only ever shows the 5 most recent entries, so the safety limit's original trip — a little further back in the timeline — has already scrolled out of view here; the next screen shows it in the full downloadable file. The broader, workspace-wide stop shown earlier in this walkthrough correctly does not appear on any single case's history at all — that's tracked separately, in the screen after next.
The same case's complete exportable history, downloaded for real and opened here — scrolled to the exact entry recording when this case's automatic safety limit tripped. The downloaded file also carries the matching reset entry moments later. This is what an auditor actually receives: not a five-item preview, but the full, timestamped file.
Every completed test run can be exported as a standalone certificate carrying a cryptographic fingerprint of its own contents — change a single number in this document after the fact, and the fingerprint no longer matches. This particular certificate is candid about its own scope, too: it carries its own visible warning that this run used deterministic checks only, with no AI judge configured for it, and should not be treated as full quality-assurance evidence on its own. A trust document that flags its own limits is more credible than one that never does.
Six real AI model connections, each one configured with the customer's own account and credentials — never a value we can see once it's saved, only a status and a connection test result. We run entirely on the customer's own AI spend; we never see or bill for the underlying usage.
A control to automatically hash sensitive field values — Social Security numbers, dates of birth, medical details — before they're ever written to the database. Shown here exactly as configured today: off, which is the safe default until an administrator deliberately turns it on.
This particular workspace is a development environment with no license file installed, and the product says so directly rather than inventing a plan name. A production deployment installs a signed license through this same screen, which then shows real usage against real limits.