Every Flag Justified: The Complete DeepSeek-V4.1-Flash Lane Recipe for Two RTX PRO 6000

A copy-paste serving recipe for DeepSeek-V4.1-Flash at EXL3 2bpw on 2× RTX PRO 6000 Blackwell (sm_120): rootless podman over the diffbot pack, eager mode forced by the Engram-on-NVMe path, the generation_config.json that ships official sampling to every client, the reasoning-effort knob that wants strings and not integers, and the measured performance envelope of the result.

September 12, 2026

The Word Salad Only opencode Could See: Hunting a Phantom Garbage Bug on DeepSeek-V4.1-Flash

CLI-dependent degradation on a quantized 552B MoE: concurrency hammers, a logging proxy that exposed what opencode really sends (nothing — and that was the bug), one reproduced episode of fluent Chinese nonsense in twelve cycles, and the missing generation_config.json that separated an uncapped top_p 1.0 tail from DeepSeek’s official recipe. Plus the reasoning-effort knob that exists as 0-100 in one engine and string levels in another.

September 12, 2026

552B Parameters and the 203 GB You Keep on Disk: DeepSeek-V4.1-Flash on Two RTX PRO 6000

The quant landscape sweep that ended at a 2.0bpw EXL3 pack built for exactly two RTX PRO 6000s, why 196B of Engram tables live on NVMe in every configuration that exists, what eager-mode serving actually measures (94-104 tok/s decode, 2.7K tok/s prefill), and the AM5 ceiling that keeps the last 15% of the recipe’s performance behind a RAM upgrade the platform cannot physically provide.

September 12, 2026