Stratt Labs — Methodology

How we test.

Real agents, real journeys, real surfaces — with verdicts engineered to be distrusted-by-default.

Methodology v1 · published 2026-07-08 · if the corroborator roster or any load-bearing rule below changes, this page changes the same day · also published as markdown

01 · Surfaces

Real agents, real journeys, real surfaces

We point production AI agents — the same model families consumers use — at live websites and ask them to complete a real journey: find a product and reach checkout, choose dates and reach the booking form, sign up, or make contact. No emulators, no synthetic DOM, no headless shortcuts a real agent wouldn't take. What we observe is what an agent-assisted customer experiences.

02 · Corroboration

Two model families, or it isn't a finding

A single agent's failure can be the agent's fault. Before we call anything a finding about a site, the same failure must reproduce across two independent model families (currently OpenAI and Anthropic), multiple runs each, stalling at the same step for the same class of reason. One family's stumble is noted honestly as inconclusive — not dressed up as a result.

03 · Calibration

Calibration before claims

The agents are tested against a fixed set of reference journeys with known outcomes before their runs count. An agent that can't pass the reference set that day doesn't get to generate findings that day.

04 · Attribution

Every run gets an honest cause, not a convenient one

Each corroborated outcome is attributed to exactly one class — and the burden of proof sits on the most serious claim:

  • SITE-DEFECT the surface demonstrably blocks a competent agent. Highest evidence bar; human-verified before it leaves the building.
  • AGENT-GAP the surface is fine; today's agents aren't. Yes, we publish this against our own thesis — it is the most common honest outcome.
  • ENV-TRANSIENT our side, the network, or quota. Never attributed to the site.
  • INCONCLUSIVE the evidence doesn't decide. We say so instead of rounding.
05 · Evidence

What counts as evidence

Every claim traces to persisted run records: step-by-step logs, screenshots, network events, and reproduction counts. Findings carry the exact step, the observed mechanism, and the evidence trail. Reports are generated from the run records — never written first and evidenced later. Sealed receipts are hash-chained so they can be verified, not just believed.

06 · Grades

The grades

Journeys earn an Agent-Readiness grade from evidence, not judgment:

  • Atwo families complete cleanly.
  • Bcompletes, with friction or one family.
  • Ccompletions and stalls mixed: a coin flip.
  • Dno run completes and two families corroborate the same stall.
  • Fa verified wall makes the site invisible to agents.
  • UNGRADEDthe evidence doesn't support a grade.
07 · Pre-registration

Pre-registered studies

When we study a market segment, the study is pre-registered: the sample, journey definitions, model families, run counts, and grading rules are frozen before the first run — and committed to cryptographically, so nobody (including us) can quietly move the goalposts afterwards. Study rules:

  1. The sample is chosen from public market lists before contact, never hand-picked for failures.
  2. Published results are aggregate-only — completion rates and failure classes across the sample. No participant is named or identifiable in anything public.
  3. Every participant sees their own results first — full private run evidence — with a 10-business-day right of reply before any finding is treated as final. If a reply shows our instrument erred, the finding is withdrawn in writing, and the withdrawal is part of the record.
  4. No bookings, orders, accounts, or submissions are ever completed. Runs hard-stop before any legally binding step (in DACH, the "zahlungspflichtig bestellen / verbindlich buchen" boundary) — every run, enforced in code.
  5. Results are a dated snapshot. Agents improve monthly; a stall observed in July is a fact about July. We date every claim and never present old failures as current ones.
current pre-registered study

European internet booking engines — sample of 18 engines, frozen 2026-07-08, before any participant contact. The full list is disclosed privately to each participant; publicly we commit to it by hash, so participants stay unnamed while the freeze stays verifiable.

sha-256 · frozen sample file f43b8a843f3a2c5f416256fa4b30504227511a03d9f02b95f396228143f41430

08 · Boundaries

What we never do

  • We never place orders, complete bookings, create accounts, or submit forms on production systems.
  • We never fabricate or extrapolate. If our tooling errs, we withdraw the finding and say why — in writing, on the record. INCONCLUSIVE is a verdict we publish, not a failure we hide.
  • We pace politely, identify honestly, and audit only public surfaces.
09 · Independence

Independence

We are not your CDN, not your payment provider, not your platform. We never sell fixes or remediation — verification is the entire product, which is the only reason our receipts are worth anything. We do offer continuous verification (re-running journeys on every release, with regression alerts); study participation and study results are free and are never conditioned on any commercial relationship, and a participant's commercial status never influences a verdict — the run records make that checkable.

Questions about the methodology or a study: audit@strattlabs.com