Skip to content
MetaBench
Model costsMethodologyCache labContact
MetaBench / Methodology

How the cache lab works

One variable at a time

Each live run gives the prefix a unique experiment marker, then sends an uncached baseline, a cache-enabled first pass and an identical replay sequentially. The changed-prefix scenario adds a fourth request that changes text before the breakpoint. Short-prefix mode limits context to 200 characters.

Measurements and interpretation

The report preserves input_tokens, output_tokens, cache_creation_input_tokens and cache_read_input_tokens from Claude. A positive read count confirms reuse. A write without a read can have several causes; the report does not label every write a proven expiry or invalidation. Requests are timed end to end.

Model selection and pricing

The lab supports Sonnet 4.6, Sonnet 5.5, Opus 5.5 and Haiku 4.5. A shared verified table supplies request prices and model-specific cache minimums. Sonnet 4.6 remains the default (1,024-token minimum); Sonnet 5.5 and Opus 5.5 require 512 tokens, and Haiku 4.5 requires 4,096. The live character-based prefix estimate warns about short inputs but does not replace Claude tokenization. Unknown models are rejected before API calls.

Costs combine uncached input, five-minute/one-hour writes, reads and output at their respective rates. Compare the current standard price table. Prices were checked October 10, 2026; estimates exclude taxes, account discounts and residency modifiers.

Anthropic prompt caching documentation

What is ready

Sample experiments, a server-side live request path, token accounting, diagnostics and JSON/CSV exports are implemented. Live Claude Sonnet 4.6 calls have been validated on our Cloudflare Worker. Email us for a token to run live requests. Sample mode uses prepared results. Live tests so far cover Sonnet 4.6; the other listed models have not yet been tested here.

Measured API validation — October 11, 2026

One controlled Cloudflare Workers run used Claude Sonnet 4.6, a reusable service-policy context and a short summary question. Each cache breakpoint had a five-minute lifetime. These are actual API usage counters, not prepared demo values.

RequestUncached inputCache writeCache readInput estimate
Baseline3,59800$0.010794
First cache pass193,5790$0.013478
Identical replay1903,579$0.001131
Changed prefix193,6050$0.013576

The four requests cost an estimated $0.041364 including output. A separate three-request short-prefix test produced no cache writes or reads ($0.001272 estimated). This checks the integration and controlled scenarios; it does not establish a production savings rate or isolate every cause of a cache miss.

Download test results (JSON) · Request live access

Open-source foundation

MetaBench uses the MIT-licensed codebase from nMaroulis/metabench. The frontend uses its React, Tailwind and theme components. The current version introduces Cache Lab for Claude request replay and token-cost diagnostics. The leaderboard is not available in this version.

© 2026 MetaBench
MethodologyPrivacyTermsContact