DeepSeek-V2-Lite-Chat-4bit-mlx
quality budget ≤ +0.5% perplexity · live optimization
+0.2%
throughput, same memory
42.33
tok/s at budget 16 (2.02 GB)
warm @ 0σ
selected policy · quality 0%
Policy tournament
Every configuration measured on the real paging runtime with physical (cache-bypassed) expert reads. Residency-bias policy per arXiv:2412.00099.
| budget | resident | policy | bias | tok/s | p95 | reads/tok | quality | |
|---|---|---|---|---|---|---|---|---|
| 16 | 2.02 GB | warm | 0σ | 42.33 | 65.8 ms | 194.3 MB | — | winner |
| 16 | 2.02 GB | none | — | 42.25 | 65.8 ms | 194.3 MB | — | |
| 16 | 2.02 GB | hybrid | 0σ | 41.45 | 64.2 ms | 194.3 MB | — | |
| 16 | 2.02 GB | cache | 0σ | 37.58 | 72 ms | 194.3 MB | — | |
| 32 | 4.05 GB | none | — | 45.49 | 48.2 ms | 94.6 MB | — | |
| 32 | 4.05 GB | warm | 0σ | 44.99 | 46.8 ms | 94.6 MB | — | |
| 32 | 4.05 GB | cache | 0σ | 44.74 | 49.1 ms | 94.6 MB | — | |
| 32 | 4.05 GB | hybrid | 0σ | 44.27 | 49.8 ms | 94.6 MB | — |
Quality calibration
Teacher-forced perplexity cost as bias strength increases — the winner comes from the safe knee under the quality budget.
Provenance: model mlx-community/DeepSeek-V2-Lite-Chat-4bit-mlx · quality budget +0.5% · single-node MLX paging runtime. Rows flagged degenerate produced looping output and are excluded from selection.