Attention Fails Before the Guard Does: Measuring Context Management at 700K Tokens

fleet-brain’s CLM harness drove real omp sessions past the local 524K window: promotion fires at 84% of the window and rides the prompt cache; with promotion disabled, a snapcompact compaction cut 445K tokens to 45K while preserving gist and destroying every verbatim canary; the headline surprise — canary recall collapsed between 389K and 443K with no context-management event at all, meaning attention had already failed exactly where the managed boundary sits. Plus: the overflow-withhold mechanism I set out to observe is unreachable on the tool surface, because every tool result is capped upstream.

October 5, 2026
Terminal-style card: local-first AI model routing across three boxes and two cloud providers

One Coding Agent, Three GPUs, Two Cloud Keys: Local-First Routing and the Judge That Was Quietly Burning My 5090

An audit of my coding agent’s model routing turned into a routing philosophy plus three real bugs. The principle: route by token volume, not vibes — the default loop, subagent fan-out, and background work run on local GPUs for free, while planning, review, and deep reasoning buy cloud quality. Found along the way: a context-promotion target pointing at a provider I have no credential for (a silent no-op), omp’s internal judge role degrading to prompted 27B judging because no TypeSafe key exists (fixed with Ollaya’s winnow:e4b decision model on the senses node, calibrated answers in 0.93 s on CPU), and the DeepSeek V4.1 naming trap where the literal model id does not exist — deepseek-flash IS V4.1-Flash, and deepseek-v4-pro server-routes to it anyway. Plus the GLM-5.3 flash/base vision trap, verified with live image requests.

October 2, 2026