Terminal-style card: local-first AI model routing across three boxes and two cloud providers

One Coding Agent, Three GPUs, Two Cloud Keys: Local-First Routing and the Judge That Was Quietly Burning My 5090

An audit of my coding agent’s model routing turned into a routing philosophy plus three real bugs. The principle: route by token volume, not vibes — the default loop, subagent fan-out, and background work run on local GPUs for free, while planning, review, and deep reasoning buy cloud quality. Found along the way: a context-promotion target pointing at a provider I have no credential for (a silent no-op), omp’s internal judge role degrading to prompted 27B judging because no TypeSafe key exists (fixed with Ollaya’s winnow:e4b decision model on the senses node, calibrated answers in 0.93 s on CPU), and the DeepSeek V4.1 naming trap where the literal model id does not exist — deepseek-flash IS V4.1-Flash, and deepseek-v4-pro server-routes to it anyway. Plus the GLM-5.3 flash/base vision trap, verified with live image requests.

October 2, 2026

Every Flag Justified: The Complete DeepSeek-V4.1-Flash Lane Recipe for Two RTX PRO 6000

A copy-paste serving recipe for DeepSeek-V4.1-Flash at EXL3 2bpw on 2× RTX PRO 6000 Blackwell (sm_120): rootless podman over the diffbot pack, eager mode forced by the Engram-on-NVMe path, the generation_config.json that ships official sampling to every client, the reasoning-effort knob that wants strings and not integers, and the measured performance envelope of the result.

September 12, 2026

The Word Salad Only opencode Could See: Hunting a Phantom Garbage Bug on DeepSeek-V4.1-Flash

CLI-dependent degradation on a quantized 552B MoE: concurrency hammers, a logging proxy that exposed what opencode really sends (nothing — and that was the bug), one reproduced episode of fluent Chinese nonsense in twelve cycles, and the missing generation_config.json that separated an uncapped top_p 1.0 tail from DeepSeek’s official recipe. Plus the reasoning-effort knob that exists as 0-100 in one engine and string levels in another.

September 12, 2026

552B Parameters and the 203 GB You Keep on Disk: DeepSeek-V4.1-Flash on Two RTX PRO 6000

The quant landscape sweep that ended at a 2.0bpw EXL3 pack built for exactly two RTX PRO 6000s, why 196B of Engram tables live on NVMe in every configuration that exists, what eager-mode serving actually measures (94-104 tok/s decode, 2.7K tok/s prefill), and the AM5 ceiling that keeps the last 15% of the recipe’s performance behind a RAM upgrade the platform cannot physically provide.

September 12, 2026

An Uncensored 305B With Eyes: DeepSeek-V4-Flash-Vision on 2× RTX PRO 6000, and the Two-Engine Week That Tamed It

The uncensored build of DeepSeek’s Vision-Exp is a byte-for-byte drop-in — which meant every failure afterwards was mine to fix. A sparse-MLA kernel that rejects vision prefill on SM120, a streaming tool-call parser that corrupted exactly the arguments coding agents live on, and a health check that proved the port but not the model. 142 tok/s single-stream, native vision, zero refusals, one new router feature.

September 4, 2026

AIOS: a Three-Node Self-Hosted Agentic Coding Stack with DeepSeek V4 Flash, vLLM, and a Custom Model Router

Instead of picking one model per task, I made three machines cooperate: a big thinking model for primary reasoning, a fast local 27B for subagent loops and images, a router with automatic failover in front of everything, and shared agent memory. Total cost: electricity. Here is the full architecture, the benchmarks, and the lessons learned the hard way.

August 21, 2026