Estimated savings

Computed from the measured throughput difference at a fixed memory budget, using your infrastructure rate — never public list prices. Everything on this page is an estimate; verified savings require production telemetry.

Your rates

Single-node decode model: node-hours = tokens ÷ measured tok/s. Batching, prefill, and utilization effects are not yet modeled.

This month (estimated)
ESTIMATE — not verified
$190,937
baseline-equivalent cost
$94,793
optimized cost (cache @ 0.125σ)
$96,144
estimated saving (101.4% faster)
Quality cost of the deployed policy: +0.53% perplexity — inside the run's stated budget. Verified savings appear only once production telemetry confirms them.