DeepSeek-V2-Lite-Chat-4bit-mlx
quality budget ≤ +2% perplexity · live optimization
+101.4%
throughput, same memory
19.78
tok/s at budget 32 (4.05 GB)
cache @ 0.125σ
selected policy · quality +0.53%
Hard quality gate: GSM8K 7/12 at winner vs 6/12 baseline · McNemar p = 1 · passed
Policy tournament
Every configuration measured on the real paging runtime with physical (cache-bypassed) expert reads. Residency-bias policy per arXiv:2412.00099.
| budget | resident | policy | bias | tok/s | p95 | reads/tok | quality | |
|---|---|---|---|---|---|---|---|---|
| 16 | 2.02 GB | warm | 0.125σ | 8.43 ±0.87 | 179.8 ms | 378.8 MB | — | |
| 16 | 2.02 GB | hybrid | 0.125σ | 7.3 ±1 | 207.4 ms | 370.8 MB | — | |
| 16 | 2.02 GB | none | — | 7.19 ±0.3 | 202.8 ms | 397 MB | — | |
| 16 | 2.02 GB | cache | 0.125σ | 6.75 ±0.04 | 225.3 ms | 357.2 MB | — | |
| 32 | 4.05 GB | cache | 0.125σ | 19.78 ±0.76 | 84.8 ms | 168.2 MB | — | winner |
| 32 | 4.05 GB | hybrid | 0.125σ | 18.02 ±1 | 98.6 ms | 169.6 MB | — | |
| 32 | 4.05 GB | warm | 0.125σ | 13.83 ±3.73 | 127.1 ms | 176.4 MB | — | |
| 32 | 4.05 GB | none | — | 9.82 ±1.02 | 167.9 ms | 212.1 MB | — |
Quality calibration
Teacher-forced perplexity cost as bias strength increases — the winner comes from the safe knee under the quality budget.
Provenance: model mlx-community/DeepSeek-V2-Lite-Chat-4bit-mlx · quality budget +2% · single-node MLX paging runtime. Rows flagged degenerate produced looping output and are excluded from selection.