DeepSeek-V2-Lite-Chat-4bit-mlx

quality budget ≤ +2% perplexity · live optimization

+11.6%
throughput, same memory
22.34
tok/s at budget 32 (4.05 GB)
warm @ 0.125σ
selected policy · quality +1.1%
standard20.01hybrid @ 0.125σ22.16cache @ 0.125σ22.33warm @ 0.125σ22.34tok/s at budget 32 · error bars are 95% CI where measured

Policy tournament

Every configuration measured on the real paging runtime with physical (cache-bypassed) expert reads. Residency-bias policy per arXiv:2412.00099.

budgetresidentpolicybiastok/sp95reads/tokquality
162.02 GBcache0.125σ13.15111.6 ms357.2 MB
162.02 GBhybrid0.125σ12.53115.2 ms370.8 MB
162.02 GBwarm0.125σ12.48113.9 ms378.8 MB
162.02 GBnone12.12120.6 ms397 MB
324.05 GBwarm0.125σ22.3473.3 ms176.4 MBwinner
324.05 GBcache0.125σ22.3374.9 ms168.2 MB
324.05 GBhybrid0.125σ22.1679.5 ms169.6 MB
324.05 GBnone20.0182.4 ms212.1 MB

Quality calibration

Teacher-forced perplexity cost as bias strength increases — the winner comes from the safe knee under the quality budget.

+0.0%+3.7%+7.4%perplexity cost vs bias strength (σ units)cachewarm

Provenance: model mlx-community/DeepSeek-V2-Lite-Chat-4bit-mlx · quality budget +2% · single-node MLX paging runtime. Rows flagged degenerate produced looping output and are excluded from selection.