DeepSeek-V2-Lite-Chat-4bit-mlx
quality budget ≤ +2% perplexity · live optimization
+11.6%
throughput, same memory
22.34
tok/s at budget 32 (4.05 GB)
warm @ 0.125σ
selected policy · quality +1.1%
Policy tournament
Every configuration measured on the real paging runtime with physical (cache-bypassed) expert reads. Residency-bias policy per arXiv:2412.00099.
| budget | resident | policy | bias | tok/s | p95 | reads/tok | quality | |
|---|---|---|---|---|---|---|---|---|
| 16 | 2.02 GB | cache | 0.125σ | 13.15 | 111.6 ms | 357.2 MB | — | |
| 16 | 2.02 GB | hybrid | 0.125σ | 12.53 | 115.2 ms | 370.8 MB | — | |
| 16 | 2.02 GB | warm | 0.125σ | 12.48 | 113.9 ms | 378.8 MB | — | |
| 16 | 2.02 GB | none | — | 12.12 | 120.6 ms | 397 MB | — | |
| 32 | 4.05 GB | warm | 0.125σ | 22.34 | 73.3 ms | 176.4 MB | — | winner |
| 32 | 4.05 GB | cache | 0.125σ | 22.33 | 74.9 ms | 168.2 MB | — | |
| 32 | 4.05 GB | hybrid | 0.125σ | 22.16 | 79.5 ms | 169.6 MB | — | |
| 32 | 4.05 GB | none | — | 20.01 | 82.4 ms | 212.1 MB | — |
Quality calibration
Teacher-forced perplexity cost as bias strength increases — the winner comes from the safe knee under the quality budget.
Provenance: model mlx-community/DeepSeek-V2-Lite-Chat-4bit-mlx · quality budget +2% · single-node MLX paging runtime. Rows flagged degenerate produced looping output and are excluded from selection.