How the cache lab works
One variable at a time
Each live run gives the prefix a unique experiment marker, then sends an uncached baseline, a cache-enabled first pass and an identical replay sequentially. The changed-prefix scenario adds a fourth request that changes text before the breakpoint. Short-prefix mode limits context to 200 characters.
Measurements and interpretation
The report preserves input_tokens, output_tokens, cache_creation_input_tokens and cache_read_input_tokens from Claude. A positive read count confirms reuse. A write without a read can have several causes; the report does not label every write a proven expiry or invalidation. Requests are timed end to end.
Model selection and pricing
The lab supports Sonnet 4.6, Sonnet 5.5, Opus 5.5 and Haiku 4.5. A shared verified table supplies request prices and model-specific cache minimums. Sonnet 4.6 remains the default (1,024-token minimum); Sonnet 5.5 and Opus 5.5 require 512 tokens, and Haiku 4.5 requires 4,096. The live character-based prefix estimate warns about short inputs but does not replace Claude tokenization. Unknown models are rejected before API calls.
Costs combine uncached input, five-minute/one-hour writes, reads and output at their respective rates. Compare the current standard price table. Prices were checked October 10, 2026; estimates exclude taxes, account discounts and residency modifiers.
Anthropic prompt caching documentation
What is ready
Sample experiments, a server-side live request path, token accounting, diagnostics and JSON/CSV exports are implemented. Live Claude Sonnet 4.6 calls have been validated on our Cloudflare Worker. Email us for a token to run live requests. Sample mode uses prepared results. Live tests so far cover Sonnet 4.6; the other listed models have not yet been tested here.
Measured API validation — October 11, 2026
One controlled Cloudflare Workers run used Claude Sonnet 4.6, a reusable service-policy context and a short summary question. Each cache breakpoint had a five-minute lifetime. These are actual API usage counters, not prepared demo values.
| Request | Uncached input | Cache write | Cache read | Input estimate |
|---|---|---|---|---|
| Baseline | 3,598 | 0 | 0 | $0.010794 |
| First cache pass | 19 | 3,579 | 0 | $0.013478 |
| Identical replay | 19 | 0 | 3,579 | $0.001131 |
| Changed prefix | 19 | 3,605 | 0 | $0.013576 |
The four requests cost an estimated $0.041364 including output. A separate three-request short-prefix test produced no cache writes or reads ($0.001272 estimated). This checks the integration and controlled scenarios; it does not establish a production savings rate or isolate every cause of a cache miss.
Download test results (JSON) · Request live access
Open-source foundation
MetaBench uses the MIT-licensed codebase from nMaroulis/metabench. The frontend uses its React, Tailwind and theme components. The current version introduces Cache Lab for Claude request replay and token-cost diagnostics. The leaderboard is not available in this version.