Terminal-style card: local-first AI model routing across three boxes and two cloud providers

One Coding Agent, Three GPUs, Two Cloud Keys: Local-First Routing and the Judge That Was Quietly Burning My 5090

An audit of my coding agent’s model routing turned into a routing philosophy plus three real bugs. The principle: route by token volume, not vibes — the default loop, subagent fan-out, and background work run on local GPUs for free, while planning, review, and deep reasoning buy cloud quality. Found along the way: a context-promotion target pointing at a provider I have no credential for (a silent no-op), omp’s internal judge role degrading to prompted 27B judging because no TypeSafe key exists (fixed with Ollaya’s winnow:e4b decision model on the senses node, calibrated answers in 0.93 s on CPU), and the DeepSeek V4.1 naming trap where the literal model id does not exist — deepseek-flash IS V4.1-Flash, and deepseek-v4-pro server-routes to it anyway. Plus the GLM-5.3 flash/base vision trap, verified with live image requests.

October 2, 2026

Every Flag Justified: The Complete DeepSeek-V4.1-Flash Lane Recipe for Two RTX PRO 6000

A copy-paste serving recipe for DeepSeek-V4.1-Flash at EXL3 2bpw on 2× RTX PRO 6000 Blackwell (sm_120): rootless podman over the diffbot pack, eager mode forced by the Engram-on-NVMe path, the generation_config.json that ships official sampling to every client, the reasoning-effort knob that wants strings and not integers, and the measured performance envelope of the result.

September 12, 2026

The Word Salad Only opencode Could See: Hunting a Phantom Garbage Bug on DeepSeek-V4.1-Flash

CLI-dependent degradation on a quantized 552B MoE: concurrency hammers, a logging proxy that exposed what opencode really sends (nothing — and that was the bug), one reproduced episode of fluent Chinese nonsense in twelve cycles, and the missing generation_config.json that separated an uncapped top_p 1.0 tail from DeepSeek’s official recipe. Plus the reasoning-effort knob that exists as 0-100 in one engine and string levels in another.

September 12, 2026

552B Parameters and the 203 GB You Keep on Disk: DeepSeek-V4.1-Flash on Two RTX PRO 6000

The quant landscape sweep that ended at a 2.0bpw EXL3 pack built for exactly two RTX PRO 6000s, why 196B of Engram tables live on NVMe in every configuration that exists, what eager-mode serving actually measures (94-104 tok/s decode, 2.7K tok/s prefill), and the AM5 ceiling that keeps the last 15% of the recipe’s performance behind a RAM upgrade the platform cannot physically provide.

September 12, 2026

The 30-Second Watchdog That Ate a Four-Hour Render: Long-Form MiniMax H3 on Two RTX PRO 6000

A 15s arcade-to-live-action video edit on H3 failed four different ways before it succeeded: cudnn graph crashes on long sequences, a 30s async-output watchdog that throws away finished renders, flashinfer kernels that require tcgen05 MMA my RTX PRO 6000s don’t have, and a finisher script that reported failure after succeeding. The winning recipe: patch the watchdog via bind-mount, CUDNN_ATTN, 540p reference, 1h45m end to end.

September 6, 2026

An Uncensored 305B With Eyes: DeepSeek-V4-Flash-Vision on 2× RTX PRO 6000, and the Two-Engine Week That Tamed It

The uncensored build of DeepSeek’s Vision-Exp is a byte-for-byte drop-in — which meant every failure afterwards was mine to fix. A sparse-MLA kernel that rejects vision prefill on SM120, a streaming tool-call parser that corrupted exactly the arguments coding agents live on, and a health check that proved the port but not the model. 142 tok/s single-stream, native vision, zero refusals, one new router feature.

September 4, 2026

From locklocklock to 262K Context: GLM-5.3-Flash on Two RTX PRO 6000

A 320B sparse-hybrid MoE serving word salad at temperature 0, a repro whose bugs outlived the bug it hunted, a per-layer bisect that cleared the KDA layers and convicted the sparse full-attention path, and an EXL3/TR3 plus B12X deployment that ends the story with verified 262K-context retrieval at 39 tok/s.

August 30, 2026

Qwen3.8-Flash-Next at 512K Context on Two RTX PRO 6000: 146 t/s, a 51B-Entry Table in Host RAM, and the FP8 KV Door That Stays Shut

Flash-Next is Qwen’s sparse-attention flagship and its official FP8 recipe assumes datacenter GB300. Adapting it to 2x96 GB workstation Blackwell took one legit trick (offloading the 51B-entry PLE n-gram table to pinned host RAM, freeing ~25 GiB per GPU for KV), one config archaeology exercise (YaRN via –hf-overrides because the build has no –rope-scaling flag), and discipline about what not to retry: fp8 KV hard-requires BF16 on QSA and fails differently on vLLM and SGLang, and 786K+ contexts OOM during compile no matter the KV math. Final stack: 512K context, 1.42M-token pool, 146.5 tok/s decode with MTP ns=3, 9-11.2K tok/s prefill, vision and tools validated. The SGLang NVFP4 day-0 lane runs too (after three SM120 gates), but serves one request at 262K slower than the vLLM lane serves two at 512K, so vLLM keeps the port.

August 27, 2026

One Server, Every Mode: MiniMax H3 Goes From Static Frames to Full Video Editing on Two RTX PRO 6000

H3 ships as two checkpoints: FL2VA (text + first/last frame) and Ref2VA (references including video). Serving both DiTs from one process took three failed hypotheses, a basename-sensitive model resolver, a 66 GB minimal partition download, and one multipart field spelled plural. Now one systemd unit serves text-to-video, keyframe animation, and video editing at 56 GiB per GPU, validated end to end.

August 27, 2026

262,144 Tokens on One RTX 5090: Qwen3.8-27B with an NVFP4 KV Cache That Actually Works

My 27B coding brain was hard-capped at ~107K context by fp8 KV and fat weights. A stranger’s gist promised a 451K-token KV pool on the same GPU via NVFP4 KV cache and a patched vLLM. I ported it, pinned it, validated it with planted-password recall tests at 235K tokens — and then watched it crash spectacularly the first time a real agent used it. The final stack: full 262,144-token context, ~125 tok/s decode, vision, MTP speculative decoding, all on one card.

August 27, 2026