Stratt Labs — Research

Model-swap token inflation, measured.

What the August 2026 Claude model swaps actually do to agent input-token bills — 24 realistic agent configurations, three languages, sealed deterministic runs, full corpus published.

Dataset v1 · measured & published 2026-08-04 · raw results hash-sealed · also published as markdown

01 · Context

Why this exists

Three dated provider events landed on agent operators in one month:

  • Anthropic retired claude-opus-4-1 on 2026-08-05. Agents pinned to it don't degrade politely — the API says no. This dataset was measured on 2026-08-04, the retirement eve — the "before" side of that swap can no longer be re-measured.
  • Claude Sonnet 5's introductory pricing ($2/$10 per MTok) ends 2026-08-31; standard pricing is $3/$15 from September 1 (provider price list, checked 2026-08-04).
  • Claude models from Opus 4.7 onward use a new tokenizer documented as producing "approximately 30% more tokens for the same text," workload-dependent.

Plenty of commentary repeats the "~30%" figure. Nobody had published measured numbers for realistic agent configurations — system prompt + tool schemas + user turns — let alone for German and Czech. So we measured.

02 · Results

Headline results

Input tokens; each cell is the median of per-bundle deltas across 8 agent configurations per language, with the min…max range.

SwapENDECZ
claude-opus-4-1 → claude-opus-5 (forced: Opus 4.1 retired 2026-08-05) +21.4% (+19.0…+23.2) +20.9% (+20.0…+22.1) +10.5% (+5.1…+12.3)
claude-sonnet-4-6 → claude-sonnet-5 +7.5% (+5.5…+9.5) +11.2% (+9.4…+12.4) +5.6% (+3.4…+7.5)
claude-opus-4-6 → claude-opus-4-8 +2.0% (−0.3…+4.0) +5.6% (+4.2…+6.7) −3.0% (−7.9…−1.5)
03 · Findings

What surprised us

Our own hypothesis was wrong, and we're publishing that. We pre-registered the expectation that Czech — diacritics, rich morphology — would inflate most under the new tokenizer. The opposite is true: Czech inflates least in every pair, and got outright cheaper in the Opus 4.6 → 4.8 swap (median −3.0%). The new tokenizer family handles Czech better than the old one did. If you operate Czech-language agents, the tokenizer change is the smallest of your migration worries.

The "~30% more expensive" shorthand overstates the Sonnet swap for real agent configs. On realistic bundles we measure +5.6% to +11.2% (median, by language) for claude-sonnet-4-6 → claude-sonnet-5 — well under the documented upper band, which is honest of the provider (they say "up to"; commentary tends to drop the qualifier).

German consistently inflates more than English in the Sonnet and Opus 4.6→4.8 pairs — a data point that matters if your clients are DACH businesses.

04 · Prices

The money axis

Token deltas are only half the bill. Per the provider's price list on the measurement date:

  • P1 (forced swap): claude-opus-4-1 was $15/$75 per MTok in/out; claude-opus-5 is $5/$25 — per-token prices fall by two-thirds on both axes. Our measured input inflation (+10.5% to +21.4%) offsets only a fraction of that: input-token bills fall roughly 60% at list prices for the same traffic. Retirement forced your hand, but on the input side it forced it downhill.
  • P2 (Sonnet): two separate effects, don't conflate them. (a) Continuing Sonnet 5 users: on September 1 the intro discount ends — +50% per token, no token change. (b) Migrating from Sonnet 4.6 after September 1 (price parity at $3/$15): the bill delta is the token delta — +5.6% to +11.2%. Migrating during the intro window is net cheaper than staying on 4.6.
  • P3: claude-opus-4-6 and claude-opus-4-8 share a price ($5/$25) — the bill delta equals the token delta, including the Czech decrease.

Boundary, stated plainly: this dataset measures the input side only. Output-token volume is workload-dependent and unmeasured here — and claude-opus-5 has thinking enabled by default, which adds output tokens a static count cannot predict. Measuring what a swap does to a running agent — cost, parameters, behavior consistency — is per-agent re-verification work, which is the paid desk, not this dataset.

05 · Method

Method

  • Corpus: 24 synthetic config bundles — 8 agent archetypes (e-commerce support, booking, voice agent, invoicing assistant, lead qualification, internal docs Q&A, ERP order desk, hospitality FAQ) × 3 languages (EN/DE/CZ). Each bundle: system prompt, 3–4 tool schemas, 5 representative user turns. No client data. The corpus is published in full below — judge its representativeness yourself.
  • Measurement: the provider's count_tokens endpoint, per (bundle × model), N=3 identical calls. All 144 measured cells were run-to-run identical (the endpoint is deterministic; N=3 documents that rather than assumes it).
  • Sealing: the raw results file is committed to by SHA-256 (below).
  • Dating: measured 2026-08-04. Tokenizers don't drift daily, but every claim on this page is a claim about that date.
06 · Verification

Parameter changes, live-verified

Documentation claims about breaking parameter changes, verified against the live API on the measurement date:

ProbeDocumented behaviorMeasured
temperature=0.7 on claude-opus-5400 (parameter removed)400 confirmed
top_p=0.9 on claude-opus-5400 (parameter removed)400 confirmed
thinking budget_tokens on claude-opus-5400 (budget_tokens removed)400 confirmed
temperature=0.7 on claude-sonnet-5400 (non-default sampling rejected)400 confirmed
temperature=0.7 on claude-sonnet-4-6 (control)acceptedaccepted
plain request on claude-opus-5 (control)acceptedaccepted
07 · Data

Full distribution

No cherry-picking: every bundle, every pair, and the underlying input-token counts (median of N=3 runs per cell).

BundleLangP1 Δ%P2 Δ%P3 Δ% opus-4-1opus-5sonnet-4-6sonnet-5 opus-4-6opus-4-8
booking-agent-czcz+10.8+5.2−3.1124713821378145014311386
booking-agent-dede+20.7+10.5+5.1121314641387153213971468
booking-agent-enen+23.2+9.5+4.097912061163127411631210
docs-qa-czcz+10.4+5.3−4.8112712441246131213111248
docs-qa-dede+22.1+11.5+5.5114613991316146713301403
docs-qa-enen+21.9+6.8+0.785510421039111010391046
erp-assistant-czcz+10.0+7.3−2.8137515121473158015591516
erp-assistant-dede+21.2+11.7+6.7132416051498167315081609
erp-assistant-enen+21.5+8.7+3.4102812491212131712121253
hospitality-faq-czcz+9.6+5.6−4.2124713671359143514311371
hospitality-faq-dede+21.1+11.5+5.7123614971403156514201501
hospitality-faq-enen+20.5+7.1+1.696811661152123411521170
invoicing-assistant-czcz+10.6+6.6−2.2138315291498159715671533
invoicing-assistant-dede+20.1+12.4+6.0136016331513170115441637
invoicing-assistant-enen+21.6+9.0+3.7104512711229133912291275
lead-qual-czcz+12.2+5.6−2.5119313381331140613771342
lead-qual-dede+20.2+9.5+4.2117414111351147913581415
lead-qual-enen+20.0+6.2+0.492011041104117211041108
support-ecom-czcz+5.1+3.4−7.9127813431364141114621347
support-ecom-dede+20.0+9.4+4.4121314551392152313971459
support-ecom-enen+19.0+5.5−0.393011071114117511141111
voice-agent-czcz+12.3+7.5−1.5128114391402150714651443
voice-agent-dede+21.1+11.0+5.7124315051417157314271509
voice-agent-enen+21.4+7.9+2.396911761153124411531180
08 · Reproducibility

Reproduce it

sha-256 · raw results file 2138a6229eb5cd3600a8c7b6529d7367cb9073b35b62b03a72cd8eaf3494d82f

Method questions, corrections, or a swap you want measured: audit@strattlabs.com. We hold our own instrument to the standard we hold anyone's: measurements, never verdicts; when our instrument errs, we withdraw the finding in writing. See how we test.