DeepSeek-V2-Lite-Chat-4bit-mlx

quality budget ≤ +2% perplexity · live optimization

+101.4%
throughput, same memory
19.78
tok/s at budget 32 (4.05 GB)
cache @ 0.125σ
selected policy · quality +0.53%
Hard quality gate: GSM8K 7/12 at winner vs 6/12 baseline · McNemar p = 1 · passed
standard9.82 ±1.02warm @ 0.125σ13.83 ±3.73hybrid @ 0.125σ18.02 ±1cache @ 0.125σ19.78 ±0.76tok/s at budget 32 · error bars are 95% CI where measured

Policy tournament

Every configuration measured on the real paging runtime with physical (cache-bypassed) expert reads. Residency-bias policy per arXiv:2412.00099.

budgetresidentpolicybiastok/sp95reads/tokquality
162.02 GBwarm0.125σ8.43 ±0.87179.8 ms378.8 MB
162.02 GBhybrid0.125σ7.3 ±1207.4 ms370.8 MB
162.02 GBnone7.19 ±0.3202.8 ms397 MB
162.02 GBcache0.125σ6.75 ±0.04225.3 ms357.2 MB
324.05 GBcache0.125σ19.78 ±0.7684.8 ms168.2 MBwinner
324.05 GBhybrid0.125σ18.02 ±198.6 ms169.6 MB
324.05 GBwarm0.125σ13.83 ±3.73127.1 ms176.4 MB
324.05 GBnone9.82 ±1.02167.9 ms212.1 MB

Quality calibration

Teacher-forced perplexity cost as bias strength increases — the winner comes from the safe knee under the quality budget.

+0.0%+3.7%+7.4%perplexity cost vs bias strength (σ units)cachewarm

Provenance: model mlx-community/DeepSeek-V2-Lite-Chat-4bit-mlx · quality budget +2% · single-node MLX paging runtime. Rows flagged degenerate produced looping output and are excluded from selection.