Real agents, real journeys, real surfaces — with verdicts engineered to be distrusted-by-default.
Methodology v1 · published 2026-07-08 · if the corroborator roster or any load-bearing rule
below changes, this page changes the same day · also published as markdown
01 · Surfaces
Real agents, real journeys, real surfaces
We point production AI agents — the same model families consumers use — at live websites
and ask them to complete a real journey: find a product and reach checkout, choose dates
and reach the booking form, sign up, or make contact. No emulators, no synthetic DOM, no
headless shortcuts a real agent wouldn't take. What we observe is what an agent-assisted
customer experiences.
02 · Corroboration
Two model families, or it isn't a finding
A single agent's failure can be the agent's fault. Before we call anything a finding
about a site, the same failure must reproduce across two independent model families
(currently OpenAI and Anthropic), multiple runs each, stalling at the same step for the
same class of reason. One family's stumble is noted honestly as inconclusive — not
dressed up as a result.
03 · Calibration
Calibration before claims
The agents are tested against a fixed set of reference journeys with known outcomes
before their runs count. An agent that can't pass the reference set that day doesn't get
to generate findings that day.
04 · Attribution
Every run gets an honest cause, not a convenient one
Each corroborated outcome is attributed to exactly one class — and the burden of proof
sits on the most serious claim:
SITE-DEFECTthe surface demonstrably blocks a competent agent. Highest evidence bar;
human-verified before it leaves the building.
AGENT-GAPthe surface is fine; today's agents aren't. Yes, we publish this against our own
thesis — it is the most common honest outcome.
ENV-TRANSIENTour side, the network, or quota. Never attributed to the site.
INCONCLUSIVEthe evidence doesn't decide. We say so instead of rounding.
05 · Evidence
What counts as evidence
Every claim traces to persisted run records: step-by-step logs, screenshots, network
events, and reproduction counts. Findings carry the exact step, the observed mechanism,
and the evidence trail. Reports are generated from the run records — never written first
and evidenced later. Sealed receipts are hash-chained so they can be verified, not just
believed.
06 · Grades
The grades
Journeys earn an Agent-Readiness grade from evidence, not judgment:
Atwo families complete cleanly.
Bcompletes, with friction or one family.
Ccompletions and stalls mixed: a coin flip.
Dno run completes and two families corroborate the same stall.
Fa verified wall makes the site invisible to agents.
UNGRADEDthe evidence doesn't support a grade.
07 · Pre-registration
Pre-registered studies
When we study a market segment, the study is pre-registered: the sample, journey
definitions, model families, run counts, and grading rules are frozen before the first
run — and committed to cryptographically, so nobody (including us) can quietly move the
goalposts afterwards. Study rules:
The sample is chosen from public market lists before contact, never hand-picked
for failures.
Published results are aggregate-only — completion rates and failure classes
across the sample. No participant is named or identifiable in anything public.
Every participant sees their own results first — full private run evidence —
with a 10-business-day right of reply before any finding is treated as final. If a
reply shows our instrument erred, the finding is withdrawn in writing, and the
withdrawal is part of the record.
No bookings, orders, accounts, or submissions are ever completed. Runs
hard-stop before any legally binding step (in DACH, the "zahlungspflichtig bestellen /
verbindlich buchen" boundary) — every run, enforced in code.
Results are a dated snapshot. Agents improve monthly; a stall observed in July
is a fact about July. We date every claim and never present old failures as current
ones.
current pre-registered study
European internet booking engines — sample of 18 engines, frozen
2026-07-08, before any participant contact. The full list is disclosed
privately to each participant; publicly we commit to it by hash, so participants stay
unnamed while the freeze stays verifiable.
We never place orders, complete bookings, create accounts, or submit forms on
production systems.
We never fabricate or extrapolate. If our tooling errs, we withdraw the finding and
say why — in writing, on the record. INCONCLUSIVE is a verdict we publish, not a
failure we hide.
We pace politely, identify honestly, and audit only public surfaces.
09 · Independence
Independence
We are not your CDN, not your payment provider, not your platform. We never sell fixes or remediation — verification is the entire product, which is the only reason our receipts are worth
anything. We do offer continuous verification (re-running journeys on every release,
with regression alerts); study participation and study results are free and are never
conditioned on any commercial relationship, and a participant's commercial status never
influences a verdict — the run records make that checkable.