Estimated savings
Computed from the measured throughput difference at a fixed memory budget, using your infrastructure rate — never public list prices. Everything on this page is an estimate; verified savings require production telemetry.
Your rates
Single-node decode model: node-hours = tokens ÷ measured tok/s. Batching, prefill, and utilization effects are not yet modeled.
This month (estimated)
ESTIMATE — not verified$190,937
baseline-equivalent cost
$94,793
optimized cost (cache @ 0.125σ)
$96,144
estimated saving (101.4% faster)
Quality cost of the deployed policy: +0.53% perplexity — inside the run's stated budget. Verified savings appear only once production telemetry confirms them.