# Model-swap token inflation, measured

What the August 2026 Claude model swaps actually do to agent input-token bills — 24 realistic agent configurations, three languages, sealed deterministic runs, full corpus published.

Dataset v1 · measured & published 2026-08-04. Raw results are hash-sealed (SHA-256 `2138a6229eb5cd3600a8c7b6529d7367cb9073b35b62b03a72cd8eaf3494d82f`) and downloadable below, together with the full corpus. If we err, we withdraw the finding in writing — corrections become part of this page.

## Why this exists

Three dated provider events landed on agent operators in one month:

- **Anthropic retired claude-opus-4-1 on 2026-08-05.** Agents pinned to it don't degrade politely — the API says no. This dataset was measured on **2026-08-04, the retirement eve** — the "before" side of that swap can no longer be re-measured.
- **Claude Sonnet 5's introductory pricing ($2/$10 per MTok) ends 2026-08-31**; standard pricing is $3/$15 from September 1 (provider price list, checked 2026-08-04).
- **Claude models from Opus 4.7 onward use a new tokenizer** that the provider documents as producing "approximately 30% more tokens for the same text," workload-dependent.

Plenty of commentary repeats the "~30%" figure. Nobody had published measured numbers for realistic *agent* configurations — system prompt + tool schemas + user turns — let alone for German and Czech. So we measured.

## Headline results (input tokens, median per language)

Each cell: median of per-bundle deltas across 8 agent configurations per language, with the min…max range.

| Swap | EN | DE | CZ |
|---|---|---|---|
| **P1** claude-opus-4-1 → claude-opus-5 (the forced Aug 5 migration) | **+21.4%** (+19.0…+23.2) | **+20.9%** (+20.0…+22.1) | **+10.5%** (+5.1…+12.3) |
| **P2** claude-sonnet-4-6 → claude-sonnet-5 | **+7.5%** (+5.5…+9.5) | **+11.2%** (+9.4…+12.4) | **+5.6%** (+3.4…+7.5) |
| **P3** claude-opus-4-6 → claude-opus-4-8 | **+2.0%** (−0.3…+4.0) | **+5.6%** (+4.2…+6.7) | **−3.0%** (−7.9…−1.5) |

## What surprised us

**Our own hypothesis was wrong, and we're publishing that.** We pre-registered the expectation that Czech — diacritics, rich morphology — would inflate *most* under the new tokenizer. The opposite is true: **Czech inflates least in every pair, and got outright cheaper in the Opus 4.6 → 4.8 swap** (median −3.0%). The new tokenizer family handles Czech better than the old one did. If you operate Czech-language agents, the tokenizer change is the smallest of your migration worries.

**The "~30% more expensive" shorthand overstates the Sonnet swap for real agent configs.** On realistic bundles we measure **+5.6% to +11.2%** (median, by language) for claude-sonnet-4-6 → claude-sonnet-5 — well under the documented upper band, which is honest of the provider (they say "up to", commentary tends to drop the qualifier).

**German consistently inflates more than English** in the Sonnet and Opus 4.6→4.8 pairs — a data point that matters if your clients are DACH businesses.

## The money axis (list prices verified 2026-08-04)

Token deltas are only half the bill. Per the provider's price list on the measurement date:

- **P1 (forced swap):** claude-opus-4-1 was **$15/$75** per MTok in/out; claude-opus-5 is **$5/$25** — per-token prices fall by two-thirds on both axes. Our measured input inflation (+10.5% to +21.4%) offsets only a fraction of that: **input-token bills fall roughly 60% at list prices** for the same traffic. Retirement forced your hand, but on the input side it forced it downhill.
- **P2 (Sonnet):** two separate effects, don't conflate them. (a) *Continuing* Sonnet 5 users: on September 1 the intro discount ends — **+50% per token, no token change**. (b) *Migrating* from Sonnet 4.6 after September 1 (price parity at $3/$15): the bill delta is the token delta — **+5.6% to +11.2%**. Migrating *during* the intro window is net cheaper than staying on 4.6.
- **P3:** claude-opus-4-6 and claude-opus-4-8 share a price ($5/$25) — the bill delta equals the token delta, including the Czech *decrease*.

**Boundary, stated plainly: this dataset measures the input side only.** Output-token volume is workload-dependent and unmeasured here — and claude-opus-5 has thinking enabled by default, which adds output tokens a static count cannot predict. Measuring what a swap does to a *running* agent — cost, parameters, behavior consistency — is per-agent re-verification work, which is the paid desk, not this dataset.

## Method

- **Corpus:** 24 synthetic config bundles — 8 agent archetypes (e-commerce support, booking, voice agent, invoicing assistant, lead qualification, internal docs Q&A, ERP order desk, hospitality FAQ) × 3 languages (EN/DE/CZ). Each bundle: system prompt, 3–4 tool schemas, 5 representative user turns. No client data. The corpus is published in full below — judge its representativeness yourself.
- **Measurement:** the provider's `count_tokens` endpoint, per (bundle × model), **N=3 identical calls**. All 144 measured cells were run-to-run identical (the endpoint is deterministic; N=3 documents that rather than assumes it). Models: claude-opus-4-1, claude-opus-4-6, claude-opus-4-8, claude-opus-5, claude-sonnet-4-6, claude-sonnet-5.
- **Sealing:** the raw results file is committed to by SHA-256 (above). The measurement script is included in the raw download's metadata trail.
- **Dating:** measured 2026-08-04. Tokenizers don't drift daily, but every claim on this page is a claim about that date.

## Parameter changes, live-verified

Documentation claims about breaking parameter changes, verified against the live API on the measurement date:

| Probe | Documented behavior | Measured |
|---|---|---|
| temperature=0.7 on claude-opus-5 | 400 (parameter removed) | 400 confirmed |
| top_p=0.9 on claude-opus-5 | 400 (parameter removed) | 400 confirmed |
| thinking budget_tokens on claude-opus-5 | 400 (budget_tokens removed) | 400 confirmed |
| temperature=0.7 on claude-sonnet-5 | 400 (non-default sampling rejected) | 400 confirmed |
| temperature=0.7 on claude-sonnet-4-6 (control) | accepted | accepted |
| plain request on claude-opus-5 (control) | accepted | accepted |

## Full distribution (per-bundle medians)

No cherry-picking: every bundle, every pair, and the underlying input-token counts.

| Bundle | Lang | P1 Δ% | P2 Δ% | P3 Δ% | opus-4-1 | opus-5 | sonnet-4-6 | sonnet-5 | opus-4-6 | opus-4-8 |
|---|---|---|---|---|---|---|---|---|---|---|
| booking-agent-cz | cz | +10.8 | +5.2 | −3.1 | 1247 | 1382 | 1378 | 1450 | 1431 | 1386 |
| booking-agent-de | de | +20.7 | +10.5 | +5.1 | 1213 | 1464 | 1387 | 1532 | 1397 | 1468 |
| booking-agent-en | en | +23.2 | +9.5 | +4.0 | 979 | 1206 | 1163 | 1274 | 1163 | 1210 |
| docs-qa-cz | cz | +10.4 | +5.3 | −4.8 | 1127 | 1244 | 1246 | 1312 | 1311 | 1248 |
| docs-qa-de | de | +22.1 | +11.5 | +5.5 | 1146 | 1399 | 1316 | 1467 | 1330 | 1403 |
| docs-qa-en | en | +21.9 | +6.8 | +0.7 | 855 | 1042 | 1039 | 1110 | 1039 | 1046 |
| erp-assistant-cz | cz | +10.0 | +7.3 | −2.8 | 1375 | 1512 | 1473 | 1580 | 1559 | 1516 |
| erp-assistant-de | de | +21.2 | +11.7 | +6.7 | 1324 | 1605 | 1498 | 1673 | 1508 | 1609 |
| erp-assistant-en | en | +21.5 | +8.7 | +3.4 | 1028 | 1249 | 1212 | 1317 | 1212 | 1253 |
| hospitality-faq-cz | cz | +9.6 | +5.6 | −4.2 | 1247 | 1367 | 1359 | 1435 | 1431 | 1371 |
| hospitality-faq-de | de | +21.1 | +11.5 | +5.7 | 1236 | 1497 | 1403 | 1565 | 1420 | 1501 |
| hospitality-faq-en | en | +20.5 | +7.1 | +1.6 | 968 | 1166 | 1152 | 1234 | 1152 | 1170 |
| invoicing-assistant-cz | cz | +10.6 | +6.6 | −2.2 | 1383 | 1529 | 1498 | 1597 | 1567 | 1533 |
| invoicing-assistant-de | de | +20.1 | +12.4 | +6.0 | 1360 | 1633 | 1513 | 1701 | 1544 | 1637 |
| invoicing-assistant-en | en | +21.6 | +9.0 | +3.7 | 1045 | 1271 | 1229 | 1339 | 1229 | 1275 |
| lead-qual-cz | cz | +12.2 | +5.6 | −2.5 | 1193 | 1338 | 1331 | 1406 | 1377 | 1342 |
| lead-qual-de | de | +20.2 | +9.5 | +4.2 | 1174 | 1411 | 1351 | 1479 | 1358 | 1415 |
| lead-qual-en | en | +20.0 | +6.2 | +0.4 | 920 | 1104 | 1104 | 1172 | 1104 | 1108 |
| support-ecom-cz | cz | +5.1 | +3.4 | −7.9 | 1278 | 1343 | 1364 | 1411 | 1462 | 1347 |
| support-ecom-de | de | +20.0 | +9.4 | +4.4 | 1213 | 1455 | 1392 | 1523 | 1397 | 1459 |
| support-ecom-en | en | +19.0 | +5.5 | −0.3 | 930 | 1107 | 1114 | 1175 | 1114 | 1111 |
| voice-agent-cz | cz | +12.3 | +7.5 | −1.5 | 1281 | 1439 | 1402 | 1507 | 1465 | 1443 |
| voice-agent-de | de | +21.1 | +11.0 | +5.7 | 1243 | 1505 | 1417 | 1573 | 1427 | 1509 |
| voice-agent-en | en | +21.4 | +7.9 | +2.3 | 969 | 1176 | 1153 | 1244 | 1153 | 1180 |

## Reproduce it

- [Corpus — all 24 config bundles (JSON)](/research/model-swap-corpus-v1.json)
- [Raw sealed results (JSON)](/research/model-swap-results-raw-2026-08-04.json) — SHA-256 `2138a6229eb5cd3600a8c7b6529d7367cb9073b35b62b03a72cd8eaf3494d82f`
- Measure your own config instead of the corpus estimate: [the free Model-Swap Exposure Check](/exposure-check)
- Method questions, corrections, or a swap you want measured: [audit@strattlabs.com](mailto:audit@strattlabs.com)

We hold our own instrument to the standard we hold anyone's: measurements, never verdicts; when our instrument errs, we withdraw the finding in writing. See [how we test](/methodology).
