DeepSeek-V2-Lite-Chat-4bit-mlx

quality budget ≤ +0.5% perplexity · live optimization

+0.2%
throughput, same memory
42.33
tok/s at budget 16 (2.02 GB)
warm @ 0σ
selected policy · quality 0%
cache @ 0σ37.58hybrid @ 0σ41.45standard42.25warm @ 0σ42.33tok/s at budget 16 · error bars are 95% CI where measured

Policy tournament

Every configuration measured on the real paging runtime with physical (cache-bypassed) expert reads. Residency-bias policy per arXiv:2412.00099.

budgetresidentpolicybiastok/sp95reads/tokquality
162.02 GBwarm42.3365.8 ms194.3 MBwinner
162.02 GBnone42.2565.8 ms194.3 MB
162.02 GBhybrid41.4564.2 ms194.3 MB
162.02 GBcache37.5872 ms194.3 MB
324.05 GBnone45.4948.2 ms94.6 MB
324.05 GBwarm44.9946.8 ms94.6 MB
324.05 GBcache44.7449.1 ms94.6 MB
324.05 GBhybrid44.2749.8 ms94.6 MB

Quality calibration

Teacher-forced perplexity cost as bias strength increases — the winner comes from the safe knee under the quality budget.

+0.0%+3.7%+7.4%perplexity cost vs bias strength (σ units)cachewarm

Provenance: model mlx-community/DeepSeek-V2-Lite-Chat-4bit-mlx · quality budget +0.5% · single-node MLX paging runtime. Rows flagged degenerate produced looping output and are excluded from selection.